GPT-5 over-explores by default in agentic work, and AI Central's GPT-5 Prompting Guide treats that as a dial you set rather than a flaw you tolerate. Lower the reasoning effort and add explicit stop criteria when you want speed. Raise it and add a persistence block when you want the model to finish without checking in. The guide also covers self-written rubrics, per-tool uncertainty thresholds, and why GPT-5 in the API answers in plain text.
The dial almost nobody sets on purpose
Most prompt advice treats model behaviour as fixed and the prompt as a polite request. This guide starts from the opposite position. GPT-5's exploration budget is configurable, and its factory setting is thoroughness, which is sensible by default and wrong for a workflow that runs a few hundred times a day.
GPT-5 is, by default, thorough and comprehensive when trying to gather context in an agentic environment to ensure it will produce a correct answer.
That line, from AI Central's GPT-5 Prompting Guide, is the premise for the rest of the document. Thoroughness is paid for in tool calls and in seconds spent watching a spinner. The guide gives two levers, the reasoning effort parameter, which it says many workflows run at medium or even low with consistent results, and criteria in the prompt telling the model how far to range.
Stop criteria are the hard part
The first snippet is labelled context gathering, and its shape matters more than its wording. A goal, get enough context fast and parallelise discovery. A method, start broad, fan out to focused subqueries, read the top hits, avoid over-searching. Then the part worth stealing, a section called early stop criteria naming two conditions. You can name the exact content to change. Or the top hits converge, roughly seventy percent, on one area.
Telling a model what to do is easy. Telling it when it has enough is the hard instruction, and most prompt libraries never try. A rough convergence threshold turns a judgement call into something the model checks mid-task without asking you. The block also caps retries with an escalate once rule, one refined batch if signals conflict, then proceed. That word, once, is what stops a research loop recursing until the budget is gone.
When you want the opposite behaviour
The second section inverts it. For maximum autonomy, the guide raises reasoning effort and adds a persistence block telling the model it is an agent, that it keeps going until the query is completely answered, ends its turn only when sure the problem is solved, and does not hand back on uncertainty.
That is not a politeness setting. It is a decision about who owns ambiguity, and a persistence prompt says the model owns it alone. What that buys when it goes wrong is confident wrong work, fast. The guide qualifies it immediately, and the qualification is the most transferable idea in the document.
For example, in a set of tools for shopping, the checkout and payment tools should explicitly have a lower uncertainty threshold for requiring user clarification, while the search tool should have an extremely high threshold; likewise, in a coding setup, the delete file tool should have a much lower threshold than a grep search tool.
AI Central's guide offers that as an example, but the rule underneath generalises. The threshold for stopping to ask a human should track how expensive an action is to undo, not how uncertain the model feels. Payment and deletion are irreversible, search and grep cost nothing. Ship a persistence prompt with a per-tool threshold map, because one blanket instruction gives the delete button the same courage as the search box.
Make the model write the standard first
For zero-to-one app generation, the guide reports that quality improves when the model executes against a rubric it builds itself. Its self-reflection snippet runs in three beats. Think of a rubric until you are confident. Think deeply about what makes a world-class result. Then use that rubric to iterate internally before answering.
The reason it works is unglamorous. World-class is an adjective, not an instruction, and no model can check its work against an adjective. Naming criteria first makes the target testable, and the test runs inside the model's own turn rather than in a review cycle after you have read a mediocre draft.
For edits to code that already exists, the guide notes GPT-5 hunts for reference context unprompted, reading package.json to see what is installed, and that naming the codebase's engineering principles, directory structure and best practices sharpens the fit.
Less reasoning means more prompt, not less
The section on system prompt reminders carries the counterintuitive finding. Prompt sensitivity is highest, not lowest, at minimal reasoning, where the guide says performance varies more drastically with the prompt than at higher levels. Four moves matter most there. Summarise the thought process at the start of the answer. Ask for tool-calling preambles that update the user. Disambiguate tool instructions and add persistence reminders. Plan explicitly in the prompt.
Prompting the model to give a brief explanation summarizing its thought process at the start of the final answer, for example via a bullet point list, improves performance on tasks requiring higher intelligence.
That is the first of the guide's four points, and what joins them is that planning has to live somewhere. At high reasoning effort the model plans privately, in tokens you never see and still pay for. At minimal reasoning those tokens are scarce, so the plan comes from you instead. Turn reasoning down to control cost and your prompts should get longer, the reverse of what most teams do when they trim.
Why your API output arrives as flat text
By default, GPT-5 in the API does not format its final answers in Markdown, in order to preserve maximum compatibility with developers whose applications may not support Markdown rendering.
AI Central's guide fixes that in two lines, telling the model to use Markdown only where semantically correct and to use backticks when it does. More useful is what it reports next. Adherence to Markdown instructions in the system prompt can degrade over a long conversation, and appending the instruction again every three to five user messages restored consistent adherence.
Read that as a general property, not a Markdown quirk. A system prompt is not a permanent setting, it is one voice competing with a conversation growing around it, and formatting is the first casualty you can see. Assume anything cosmetic decays with context length, and re-assert it on a schedule instead of debugging it as a bug.
What to do with this
Five moves, in payoff order.
- Classify each workflow as speed-first or completion-first, set reasoning effort to match, then use the context gathering block or the persistence block. Both at once gives contradictory orders.
- Write stop criteria the task can verify, a named file, a converged set of results, a passing test, not the phrase when you have enough.
- Rank every tool by how hard its action is to undo, and give the irreversible ones a low threshold for stopping to ask. Payment, deletion, anything that sends a message.
- For one-shot generation, ask for the rubric and the artefact in the same turn, rubric first. An internal review pass for a few hundred tokens.
- If you run at minimal reasoning to save money, add planning instructions and tool preambles before cutting anything else, and repeat formatting rules every few messages.
One note on sourcing. Every claim here comes from one document, the GPT-5 Prompting Guide in the AI Central library, reviewed in August 2026. Its snippets are scaffolds, not finished artefacts, and the guide says so itself when it tells you to rewrite its rules to your own taste.
Does turning reasoning effort down make GPT-5 worse?
Not automatically. The guide's position is that many workflows run at medium or even low effort with consistent results, and what you give up is exploration depth. The cost lands elsewhere. At lower effort the prompt carries the planning the model would otherwise do internally, so a short prompt plus low effort is what degrades quality.
Can I use the persistence prompt and the context gathering prompt together?
They pull in opposite directions by design. Context gathering cuts exploration to reach an answer faster, persistence raises autonomy to keep the model working. Pick the one that matches the job, and match the parameter to it, lower reasoning effort for the first and higher for the second.
Do these prompts work in ChatGPT or only through the API?
The prompt blocks are plain text and paste into any GPT-5 surface. The parameter advice is API territory, since reasoning effort is passed with the request, and the Markdown default described is specifically GPT-5 in the API. Treat the parameters as developer-facing and the prompt structures as portable.