
Anyone using AI coding agents like OpenAI’s Codex in their daily workflow knows the pattern: instruction files that tell the model how to behave keep growing over months. With the new flagship model GPT-6 Astra, that well-intentioned effort is now partly backfiring. OpenAI developer Eric Provencher warns in a guide published on September 11 that the same bloated instructions teams used to keep weaker models in line are now noticeably slowing the more capable Astra down.
Key takeaways
- OpenAI recommends developers fundamentally rework their skill descriptions, AGENTS.md files, and task prompts for GPT-6 Astra rather than simply carrying them over unchanged.
- Too many or overly long skill descriptions get automatically truncated by Codex when several are loaded at once, leading the model to make worse selection decisions.
- Astra runs tests and checks its own work on its own initiative; old instructions that explicitly demand this now just waste context space.
- Contradictory rules across different files unsettle Astra more than earlier models — it can stop mid-task or misprioritize as a result.
- OpenAI recommends narrowly scoped skill descriptions, conditional rather than blanket reading requirements, and an explicit definition of when a task counts as “done.”
The problem: rules built for the old model hold back the new one
According to Provencher, the core issue is that many teams never fundamentally cleaned up their steering files — they just kept adding to them, one model generation after another, for a full year. That made sense for weaker predecessor models: they needed explicit reminders to run tests, check their own work, or read certain documents before every change. GPT-6 Astra now does much of that on its own. Anyone who leaves the old instructions untouched isn’t just wasting context space — superfluous, sometimes contradictory directives can actually make the model perform worse than if it had no instructions on the topic at all.
An example from OpenAI’s guide illustrates this well: a skill description like “Create and validate Postgres schema migrations. Use when working with databases, queries, models, or persistence” is worded so broadly that it gets loaded for nearly every task — even ones that have nothing to do with an actual migration. The recommended alternative narrows the scope concretely: only when adding, changing, or reviewing a migration’s rollout. That may sound like a minor detail, but it has direct consequences: Codex automatically truncates descriptions once several skills are loaded at once, and the vaguer the description, the more likely the model is to pick the wrong skill or load unnecessary baggage.
AGENTS.md: mandatory reading becomes a context trap
A second common pattern involves the AGENTS.md file, where teams store core project rules for the agent. Many such files demand, across the board, that several documents be read before every change — architecture, database, and deployment docs, for instance. That might make sense for a complex feature change, but it wastes context window on a single typo fix — context the model then lacks elsewhere. OpenAI instead recommends a tiered approach: a lean top-level guide that only points to more detailed sub-documents when actually needed, instead of mandating everything upfront regardless of the task.
A third, trickier problem is one Provencher emphasizes in particular: Astra reacts more sensitively to contradictions between different instruction sources than its predecessors did. If a skill file and AGENTS.md disagree even in nuance, the model can pause mid-task, change direction, or follow a rule the user never actually intended. The advice, therefore, is not to paper over outdated or overlapping directives with yet more instructions, but to consistently remove them and clearly establish which file takes precedence when in doubt.
Why this is more than a niche problem
For solo developers, this might sound like fine-tuning. For teams with shared skill libraries, it has real consequences. Such skills are often used by several models in parallel — alongside Astra, also by older model versions still in active use. Tailoring an instruction radically to Astra’s capabilities can end up holding back other models running in the same environment. OpenAI therefore recommends not simply stripping guardrails away, but reworking them deliberately: making permissions narrower but clearer, for instance with explicit approval for safe local test runs, rather than blanket bans or blanket allowances.
The timing of the recommendation fits a pattern kabel-salat.info has observed repeatedly in recent weeks: since GPT-6 Astra launched in early September, reports of unusually high demand and capacity strain at OpenAI have piled up, as seen when OpenAI temporarily stopped accepting new ChatGPT Pro customers because of the rush on Astra. A model that works noticeably more efficiently thanks to cleaner instructions doesn’t just save developer time — it also saves compute capacity, a detail that’s unlikely to be a coincidence given the current demand situation.
Conclusion
The recommendation reads as unspectacular, but it hits a sore spot in agentic AI practice: many teams optimize their models more for the past than for the present, because instructions written once are rarely questioned from the ground up. OpenAI’s core message is, at bottom, simple — a more capable model needs less guidance, not more, and the developer’s job shifts from micromanaging behavior to clearly defining when a task is actually done. Teams that routinely review skills, AGENTS.md files, and prompts with every model switch stand to benefit noticeably; those that don’t risk making an objectively stronger model perform worse than its predecessor.

