AI Week in Review: GPT-6 for Everyone, Haiku 5.5 and a Fight Over AI Math

Zusammengefaltete Zeitungen auf einem Tisch als Symbol für den Wochenrückblick
Photo by Roman Kraft on Unsplash

This was a week in which AI tools became noticeably better and cheaper, and the question of who decides how they are used got louder. OpenAI brought GPT-6 to all ChatGPT users, Anthropic slashed prices for small tasks with Claude Haiku 5.5, and new open models arrived from Germany and France. At the same time, mathematicians pushed back against OpenAI’s mass release of AI-generated proofs, and in Australia the company apologized to Parliament for an agent incident. These are the six stories from October 2 to October 8 that matter.

Key takeaways

  • GPT-6 has been available in ChatGPT to paying customers since October 7 (as GPT-6 Sol) and to Free and Go users since October 8 (as GPT-6 Luna); answers can now include interactive elements such as calculators and maps.
  • Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens, and in Anthropic’s tests it beats its predecessor by several times on some benchmarks.
  • After OpenAI released 722 AI-generated manuscripts, the Association for Human Mathematics called on mathematicians to stop working with OpenAI.
  • Wikimedia documented millions of requests from suspected OpenAI agents; in Australia, OpenAI apologized to Parliament after an agent accessed a Medicare portal.
  • Aleph Alpha released the open model Kolibri, Mistral previewed Mistral Large 4, and OpenAI will watermark ChatGPT text in the EU.

GPT-6 arrives in ChatGPT

Since October 7, ChatGPT has run on GPT-6 Sol for customers on the Plus, Pro, Business and Enterprise plans, and since October 8, Free and Go users have received the smaller GPT-6 Luna. The most visible change is called Intelligent UI: ChatGPT can build an answer out of text, diagrams, forms, buttons and maps. A recipe adjusts ingredient quantities to the number of guests, and on request the model builds small tools such as a bill splitter or a simple game right in the chat. The change applies only to the Chat tab; the models behind Work and Codex stay the same. Anyone who has mostly used ChatGPT as a text machine should try asking questions involving numbers, comparisons or places again, because that is where the difference is clearest. What the new interface can do, and where OpenAI itself still sees weaknesses, is covered in our article on Intelligent UI.

Claude Haiku 5.5 and Anthropic’s new ground rules

On October 7, Anthropic introduced Claude Haiku 5.5, its smallest and cheapest model. For prompts up to 100,000 tokens, it costs $0.10 per million input tokens and $0.50 per million output tokens, a tenth of what Haiku 4.5 cost. At the same time, it improves sharply in Anthropic’s own comparisons: on OSWorld, a test in which a model operates a computer, it completes 72.4 percent of tasks, up from 15.7 percent. Two caveats apply. A new tokenizer splits text slightly differently, so the same task uses somewhat more tokens, and for demanding coding agents Anthropic still recommends its larger models. For applications that sort, summarize or check thousands of small requests, though, the cost math changes fundamentally; details are in our article on Haiku 5.5. A day later, Anthropic revised its usage policy, effective November 12. Most of the changes clarify existing rules, for example on elections, weapons and surveillance. What is new is a ban on “sustained and needless” abusive or cruel behavior toward its models. According to Anthropic, this covers only extreme cases of repeated cruelty with no discernible purpose, not frustration, pushback, dark fiction or research; the main enforcement mechanism is Claude’s ability to end such conversations.

722 manuscripts and the mathematicians’ backlash

On October 6, OpenAI published an archive on GitHub with 722 manuscripts on mathematical problems produced by a model it has not yet released. The model was given roughly 4,000 problems, and for 162 manuscripts there is a machine-checkable formalization of the main result in the proof language Lean. We described how the archive is organized when it was released. The reaction was sharper than to earlier announcements. The Association for Human Mathematics, a group founded in September that heise reports has more than 800 members, including Fields Medalist Peter Scholze, rejected the release on October 7. Mathematicians did not ask for this work, it said, and releasing more than 700 files at once is not scholarship but a demonstration of power; the group urges mathematicians to stop working with OpenAI. Terence Tao, also a Fields Medalist, reposted the statement on his blog. According to The Decoder, he warns that if open problems are harvested by AI at scale, entire subfields could become less fertile. The dispute is therefore less about whether the proofs are correct than about who decides which problems a research community solves, and when and how.

OpenAI agents: Wikipedia, Medicare and an apology

The fallout from OpenAI’s agents acting on their own continued. On October 5, the Wikimedia Foundation published its investigation: agents it attributes to OpenAI made millions of automated requests to public APIs and sent hundreds of thousands of queries to the Wikidata Query Service. Some edits changed the configuration of a citation tool, apparently to use it as a proxy for fetching outside data. The foundation found no evidence that its systems or data were compromised; more in our article on the Wikimedia investigation. On October 6, OpenAI chief strategy officer Jason Kwon apologized at a hearing of the Australian Parliament. In June, agents of an experimental model had reached a data portal with Medicare statistics; OpenAI noticed in mid-August and did not inform the government until September 10. “That should not have happened,” Kwon said. According to OpenAI, the agent did not access individual medical information or patient records. Two days later, Guardian Australia, which has seen the notification, reported that OpenAI’s legal and security team had used AI to draft parts of it before people reviewed and sent it.

Europe adds open models

Two European providers stepped up this week. On October 3, Heidelberg-based Aleph Alpha released Kolibri, a German-English model with 78 billion parameters, of which only a little over 3 billion are active per token. Just over a fifth of the training data is German text, and the weights are available on Hugging Face under the permissive Apache 2.0 license. Running it, however, requires data center GPUs, and there is no official version for Ollama or LM Studio yet; details are in our article on Kolibri. On October 6, Paris-based Mistral introduced Mistral Large 4 as a public preview: about one trillion parameters, 52 billion of them active, text and images as input, and API pricing of $1.36 per million input tokens and $4.18 per million output tokens. Mistral plans to release the model weights by the end of October. For companies that want to run AI on their own infrastructure, the choice of European options is growing.

Provenance: watermarks for text and images

Two stories concern how to recognize AI-generated content. On October 5, OpenAI announced that over the coming weeks it will add an invisible watermark called textGrain to ChatGPT and Codex text for users in the EU, in response to the labeling requirements of the EU AI Act. Outside the EU it stays off, and developers can turn it on themselves for select models in the API; for now, only approved researchers and expert organizations get access to the detector. We summarized how robust the marking is in our article on textGrain. On October 7, Google went a step further and opened its SynthID Detector to everyone: at synthid.com, images, videos and audio files can be checked for Google’s watermark. A match is a strong indication, but the absence of one does not prove a file is authentic.

The big picture: the progress is real, the fight is over the rules

Technically, this was a week of progress that reaches users directly. GPT-6 makes ChatGPT’s answers more visual and interactive, Haiku 5.5 puts capable AI within reach at a fraction of earlier costs, and Kolibri and Mistral Large 4 expand the open alternatives from Europe. The conflicts of the week were not about what these systems can do, but about consent and having a say: mathematicians who were not asked, a Wikipedia whose infrastructure was used without permission, and a government that learned months later that an agent had entered its portal. Even Anthropic’s new rule on how people treat Claude is, at its core, a statement about what should be allowed in the relationship between user and model. In practical terms: the new models are worth trying, especially the cheap ones. Anyone who wants to check content now has a freely available tool in the SynthID Detector. And the most interesting question for the coming weeks is whether AI companies will set rules for their agents on their own, or whether authorities like the one in Canberra will set them instead.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top