xAI: when the agent says “done” and nothing was done
I use Grok in agents for owners who do not write software. I recorded specific failures there: a task announced as done, a tool call declared but never run. I published them in April 2026. Since then I have worked with teams at xAI: reproducible cases, two public gists, three fixes merged into Hermes Agent.
Context
I use Grok in a fleet of agents for owners who do not write software. The failures I record come from there: a task announced as done, a tool call declared without execution, a trace too raw to serve as an evaluation.
On 21 April 2026 I published “The benchmark lie: why Grok 4.20 excels in benchmarks but fails in production”. On 6 May, Eric Jiang, of xAI, replied in public, under that article:
Awesome writeup, thank you for this. Would love your help to improve our model; DMing
Since then I have worked with teams at xAI.
What was done
A benchmark score describes a test bench. The question I ask is elsewhere: what the model does when a person hands it a task, in an agent loop, on their files.
I cut the failures into short cases, replayable, with a starting state and an expected state. Two of those packets are public.
Cases passed out of 12, by reasoning effort (Grok 4.3)
none 2.4/12; low 4.2/12; medium 4.7/12; high 4.8/12. Mean of 10 runs. Source: public gist of 7 May 2026. Cases passed out of 12, by reasoning effort (Grok 4.3)
Even at the best setting, fewer than half the cases pass.
On 11 April, a public gist describes a guard against hallucinated completion, taken from a Hermes session under Grok 4.20, on 10 April. On 7 May, another public gist replays 12 real cases, 10 times each, under Grok 4.3, varying the reasoning effort.
Other packets stayed private. I describe them by date, count and type. Never by their address, and never by their content.
| Date | Piece | Status |
|---|---|---|
| 11 April | Completion guard, session of 10 April | Public |
| 7 May | 12 real cases, 10 runs each | Public |
| 7–11 May | Benchmarks on real replays, including an instrumented tier: 5 conditions, 12 cases, 3 repeats | Private |
| 12 May | Report of 8 cases | Private |
| 12 May | 2 comparisons on a checkpoint | Private |
| 17 May | 5 reproductions of production blockages | Private |
| 12 June | 1 reproduction: the same path goes from 10 minutes to 33.9 seconds | Public, on the site |
On 12 June, “Making Grok act” describes a system-prompt fix. On one reproduction, the same path goes from 10 minutes to 33.9 seconds. One case. Not a benchmark.
Three fixes touching xAI are merged into Hermes Agent, with my name kept in the history. On 10 April, the native xAI provider, pull request 7372. On 9 May, passing reasoning effort to the xAI API, pull request 22807. On 12 September, a fix that removes a 400 error on the end-of-iterations summary call, pull request 109295. On 23 May, Teknium wrote, in public: “Thanks to Julien Grok Build v0.1 now has its appropriate 256K context length in Hermes”.
What changed
Public exchanges
21 Apr
“The Benchmark Lie”
6 May
Eric Jiang’s public reply
23 May
Teknium’s note, 256K context
8 Jul
Public thanks from Miłosz Jankiewicz
14 Jul
Grok 4.5 model card, FalseClaimBench
22 Jul
Article: I am not the author of FalseClaimBench
Public pieces only, linked in the sources.
Measured, in the sense of a public act. Eric Jiang’s reply opens the relationship: it is dated, it is public, it is quoted as written. On 8 July, Miłosz Jankiewicz thanked people in public “for all the support and feedback”. On 21 August, he replied in public to a suggestion on real-time voice: “Ok, will ship voice s2s. Give me a day”. Delivery of that voice is not verified.
The teams asked for reproductions, examined traces, and opened early access to checkpoints before release. The detail of those requests and of that access comes from private exchanges. I do not quote them, and I do not date them person by person.
The Grok 4.5 model card, of 14 July 2026, describes FalseClaimBench, an internal evaluation: compare what the agent claims it did with the true final state of the workspace. That is the same family of failure I have documented since May. On 22 July I wrote that I am not its author: the name, the setup and the evaluation stack are xAI’s. The convergence is the fact I keep. It does not prove that my cases were used in training, or that they changed a model.
Limits
I do not publish the private messages, or the content of the reports that stayed private.
None of these pieces establishes a contract, a payment, an official role, or a beta-tester title. The pull requests opened on 21 September for Grok 4.7, and the follow-up proposed on 22 September for voice transcription, are not counted as merged.
Sources
- “The Benchmark Lie”, 21 April 2026. Eric Jiang’s reply, 6 May 2026.
- Completion guard, public gist of 11 April 2026.
- Public benchmark of 7 May 2026: 12 real cases, 10 runs, reasoning effort under Grok 4.3.
- Teknium’s note, 23 May 2026. Thanks from Miłosz Jankiewicz, 8 July 2026. Reply on voice, 21 August 2026.
- “FalseClaimBench: the right to stop checking”, 22 July 2026. Grok 4.5 model card, 14 July 2026, section 2.11.
- Hermes Agent: pull request 7372, 22807, 109295.
- “Making Grok act”, 12 June 2026.