Leaderboard stories travel fast. This one rests on a single secondary outlet—not an official VulcanBench dump or an xAI confirmation checked for this draft.
On 6 October 2026, CryptoBriefing reported that xAI’s Grok 4.7 holds first, second and third place on the VulcanBench Frontier v4 leaderboard across three reasoning-effort settings. [1]
Treat everything below as secondary trade-press reporting unless and until primary leaderboard pages or xAI materials are verified independently.
The scores CryptoBriefing reports
According to CryptoBriefing, Grok 4.7 lets users pick how much reasoning effort the model spends. Each setting was scored separately:
- Extra-high: 93.15 (1st)
- High: 92.71 (2nd)
- Medium: 92.30 (3rd)
The closest rival named in that article is Claude Fable 5.1 at 91.84—so even the medium setting is reported ahead on this test. CryptoBriefing puts the gap from extra-high Grok 4.7 to Claude Fable 5.1 at 1.31 points. [1]
At extra-high effort, the same piece says Grok 4.7 passed all 23 behavioural-reconstruction tasks on the benchmark. In CryptoBriefing’s description, a behavioural-reconstruction task asks a model to rebuild software so that it behaves exactly like a reference—approximate output does not count. [1]
What VulcanBench is said to measure
CryptoBriefing describes VulcanBench Frontier v4 as evaluating models on real-world software-engineering problems, weighing functional correctness, code quality and complexity, with deterministic hidden tests so the model cannot see the answer key and the same code always gets the same grade. [1]
That characterisation comes from the secondary article, not from a primary VulcanBench document fetched for this draft.
Pricing and other claims in the same piece
CryptoBriefing says Grok 4.7 launched on 21 September 2026, is built on an extended base versus Grok 4.6, supports text and image inputs plus additional tools, and offers a 500,000-token context window. Pricing is reported unchanged from 4.6: $2 per million input tokens and $6 per million output tokens. The article also says xAI reported improvements versus 4.6 on CursorBench and Terminal-Bench. [1]
Those model and pricing details are again secondary trade-press claims about company positions.
Why label this carefully
CryptoBriefing is secondary coverage in the trade/crypto press. Primary VulcanBench results pages and xAI’s own launch materials were not checked for this explainer. Numbers, rival scores and the “all 23 tasks” claim could be wrong, out of date, or incomplete.
What we don’t know / What this does not prove
- Whether the leaderboard snapshot remains current after 6 October 2026.
- How VulcanBench samples tasks, prevents contamination, or weights code quality versus functional pass rates—beyond CryptoBriefing’s summary.
- Whether “effort levels” change latency and cost in ways that matter for production coding agents.
- A thin gap over one named rival on one coding benchmark does not prove overall superiority across models or workloads.
The Bottom Line
Secondary outlet CryptoBriefing says Grok 4.7 swept the top three VulcanBench Frontier v4 slots by effort level, with extra-high at 93.15 and a clean sweep of 23 behavioural-reconstruction tasks. Until primary leaderboard or xAI sources are confirmed, treat this as unverified trade-press scorekeeping—not a settled frontier ranking.
Sources
- CryptoBriefing (Diego Almada Lopez), “Grok 4.7 takes the top three spots on VulcanBench Frontier v4,” 6 October 2026 — https://cryptobriefing.com/grok-4-7-tops-vulcanbench-frontier-v4/

