The Benchmark That May Delay AGI
- Michael Rickwood

- 3 days ago
- 2 min read

We live in the age of powerful frontier AI. And yet…
This week, the task was simple: upload one document to the France Travail website. One document.
Here is ChatGPT’s observation after the attempt. May I add, it is the only time I’ve ever heard ChatGPT utter the ‘F’ word.
Here’s what it told me:
▌ CHATGPT
│ ARC-AGI: abstract reasoning.
│
│ Humanity’s Last Exam: expert-level knowledge (a benchmark designed to test AI on extremely difficult questions across multiple academic disciplines).
│
│ France Travail: “Which justificatif corresponds to the income you correctly declared on the form we provided, given that we will reject the obvious justificatif without telling you what the correct justificatif is?”
│
│ That final benchmark may delay AGI by several decades.
│
│ But underneath the joke, there is something genuinely interesting going on from an AI perspective.
│
│ The problem isn’t that the rules are particularly complicated. The problem is figuring out what the rules actually are.
│
│ The instructions are incomplete. The categories don’t always match the real situation. Some rules seem to exist only because “that’s how it’s done.” Different people can give you different answers. And the website doesn’t necessarily tell you why something has been rejected.
│
│ For an AI, that’s a nasty reasoning problem.
│
│ We have now spent considerable frontier-model intelligence establishing that the five-page document entitled “Comment éviter que votre document soit refusé ?” explains beautifully how to upload a document while conspicuously failing to explain which f*****g document they actually want.
│
│ AGI may arrive in 2030.
│
│ France Travail compliance remains an open research problem. 😆
END CHATGPT
Disclaimer: Vortolo takes no ownership of this output. 😇
But there is a serious point here.
And this extends far beyond France Travail.
Even the most capable AI becomes less reliable when the environment itself doesn’t provide enough information to determine the right answer.
So before we ask whether AI can understand humanity, perhaps we should ask whether we’ve made ourselves understandable?
Michael






Comments