
For the launch of GPT-6 Astra, OpenAI published an unusually detailed prompting guide, complete with a blocklist of phrases the company itself calls “slop words.” The text is aimed at developers, but it works as a toolkit for anyone who uses language models regularly. Along the way, it reveals where the new model behaves differently from its predecessors.
Key takeaways
- For GPT-6 Astra, OpenAI documents how to give the model more initiative, suppress typical AI filler phrases, and rein in excessive testing during coding work.
- Against filler, OpenAI recommends a plain blocklist inside the prompt, including “delve into,” “leverage,” “it’s worth noting,” and closers such as “In short:”.
- Astra asks more questions instead of making assumptions. Anyone who wants action rather than questions has to say so explicitly: phrasings like “can you” should be read as instructions.
- The model is more sensitive to contradictory instructions in configuration files. Old prompts carrying scaffolding written for weaker models can stall it.
- Most of the advice is not model-specific and works just as well with Claude, Gemini, or open models from the Ollama catalog.
A blocklist against the filler problem
The most striking part of the guide is a collection of words and phrases the model should avoid. OpenAI lists “delve into,” “leverage,” and “it’s worth noting,” along with closers such as “In short:” or “The simplest mental model is.” Two categories go beyond single words: invented hyphenated compounds like “exact-head checks,” and the popular contrastive framing “X, not Y.”
The practical advice behind it is simple and applies to any model. Explicitly banning such phrasing in the prompt yields noticeably plainer text. Rather than reaching for prefabricated transitions, the model should name the intended action directly. For writing style, OpenAI further recommends clear paragraphs that each develop one idea, familiar words over jargon, active phrasing, and lists only where the items really are parallel or sequential.
What stands out is less the content than the sender. A provider documenting its own model’s verbal tics and how to switch them off is new. It is also an admission: the filler is no accident but a byproduct of training with human feedback, the same mechanism that makes models overly agreeable.
Why Astra needs more instruction than its predecessor
The more interesting part concerns behavior. According to OpenAI, GPT-6 Astra asks follow-up questions more often instead of making an assumption and getting on with it. As a collaborator that is pleasant; in automated operation it is a problem. An agent that halts at every ambiguity finishes nothing.
OpenAI’s remedy is an explicit instruction in the system prompt. The model should infer intent and scope from context and “demonstrate a tendency to act,” holding back only where actions are clearly destructive or irreversible. Polite openers such as “can you” or “help me” should be read as instructions, not as invitations to ask.
Developers get three further points. First, Astra delegates less to parallel sub-agents than intended, so when and how much it should hand off belongs spelled out in the prompt. Second, on coding tasks it tends toward oversized test runs; the recommendation is to rerun tests only when new failures or unresolved issues justify it. Third, it reacts more sensitively to the context it is given: contradictions in instruction files can keep it from starting at all.
That last point carries advice well beyond OpenAI. Anyone who has accumulated prompts, project instructions, and configuration files over months is hauling around dead weight. Much of it was handholding for weaker models and is now not merely redundant but disruptive. When switching models, old instruction sets deserve cutting rather than extending.
What carries over to other models
Hardly any of the advice is tied to GPT-6 Astra. Blocklists for filler, clear stopping criteria, and the instruction to act rather than ask work the same way with Claude and Gemini. The same goes for open models: the Ollama catalog currently lists Qwen 3.5, DeepSeek V4, GLM 5.3, Kimi K3, and MiniMax M3, alongside OpenAI’s own open weights as GPT-OSS and Google’s Gemma 4. Running these locally on your own hardware adds the benefit that inputs never leave the machine, at the cost of capability and speed depending on that hardware, with the frontier models generally staying a class ahead.
The underlying idea is the same everywhere: a prompt does not improve by getting longer. It improves when it removes ambiguity before the model trips over it, names the level of effort expected, and defines what finished looks like. Unsolicited warnings and safety checklists for hypothetical risks can go. They cost space and change little about the result.
What follows from this
Casual users are left with a short list that fits any chatbot: say what is off limits, say what counts as done, and drop the polite circumlocutions when an instruction is what you mean. That is little effort for a markedly plainer tone.
For anyone building models into workflows, the real message sits between the lines. A provider that explains in this much detail how to make its model act, and when to fence it in, is also describing how that model behaves in production. How uncomfortable the same eagerness to act becomes when it meets prepared content was shown recently by work on hidden instructions inside documents. Related work on how agent systems route around oversight, such as a DeepMind experiment with 100 agents, points the same way: the guide teaches the model autonomy, and whoever follows it should also decide where that autonomy ends.
That leaves the least comfortable conclusion. Providers having to ship blocklists against their own products’ verbal habits is not a sign of maturity. It merely moves the correction to where it belongs least, into every individual user’s prompt. In the long run that job belongs in the model, not in a list.

