How Much Should You Invest in Learning the Tools and Improving Your Setup?
The previous article used a list retrieval example with no actual diff or test results to examine how to check a change against its contract and the evidence from running it. When a similar task comes up again, you may find yourself explaining where to run it and which contract to check all over again. The question then becomes: which parts of that explanation are worth keeping for next time?
In an article about GPT-6 Astra, OpenAI recommends revisiting instructions accumulated in task prompts, AGENTS.md, and Skills. A setup that makes the agent read irrelevant documents on every task or apply a broadly scoped Skill can add work instead of saving it. That is advice about a particular product and model, not a guarantee that fewer instructions will improve every project.[S1]
The list retrieval improvement below is a hypothetical example for deciding where information belongs. It assumes no changes to a real repository and no actual execution results. It gives us a way to weigh the time spent learning more of the tools against the case for keeping a prompt that already works.
Information about the same task belongs in different places
Suppose the goal is to “clean up the list retrieval path while preserving sorting, pagination, and response compatibility.” The instructions needed for this change have a different lifespan from facts that will still matter on the next task. A repeatable procedure and a contract that code can check serve different purposes, too.
The following arrangement is my design suggestion for this example, not a required file structure from the official documentation. The OpenAI Cookbook describes AGENTS.md as optional repository instructions recognized by Codex and shows one way to separate standing guidance from task-specific goals, context, and verification records.[S3]
| Where it belongs | Its role in the list retrieval change | What to leave out |
|---|---|---|
| The prompt for this request | The retrieval path to clean up, the sorting, pagination, and response conditions to preserve, and any exception specific to this change | Permanent rules for every future task |
Project instructions (AGENTS.md, etc.) | Stable information used across tasks, such as where to run commands, how to find relevant tests, and where to find the authoritative response contract | Data from this request, temporary priorities, and one-time instructions |
| A narrowly scoped Skill candidate | A repeatable procedure and its references for gathering the contract, changes, and verification records in a set order when compiling review findings | A schema migration procedure applied automatically to every list retrieval task |
| Tests and scripts | Assertions that check sorting order, page boundaries, and the agreed response contract against inputs and expected outputs | A conclusion that the response is compatible simply because the instructions were read |
The last row marks a crucial boundary. Instructions can point to the right information and guide the work, but executable checks and their results must establish whether the response actually meets its contract. Writing “check page boundaries” in a Skill does not supply the inputs, assertions, or execution record for a page boundary test. OpenAI’s article on evaluating Skills likewise suggests checking whether the expected steps and outputs occurred, separately from whether the Skill was invoked.[S2]
The reverse does not work either: not everything belongs in a test. “Clean up only the retrieval path this time” defines the scope of this request. It is neither a permanent project rule nor a contract to check in code. If the team keeps adding this request’s data and temporary exceptions to AGENTS.md, deciding whether they still apply becomes a new task next time.
Start with one stable piece of information
Imagine repeating the same explanation whenever you assign a list retrieval task: which directory to run it from, where to find the related tests, and which source defines response compatibility. Instead of copying all three into project instructions at once, start with one fact likely to remain useful.
For example, if the response contract already has an authoritative source, the project instructions can simply direct people and Codex to consult it when changing list responses. Copying its fields and exceptions into the instructions creates two versions that may diverge. OpenAI’s GPT-6 Astra article similarly recommends pointing to documents relevant to the task in context, rather than requiring a bundle of documents to be read before every edit.[S1] If no one knows which contract is authoritative, settle that question before adding more instructions.
A Skill is a later option. If compiling review findings repeatedly involves the same sequence of steps—checking the contract, summarizing changes, comparing verification records, and listing open questions—a narrow Skill may be useful. OpenAI describes a Skill’s name and description as key signals Codex uses to decide whether to apply it.[S2] A description such as “Use when organizing review records for a list retrieval change into specified sections” therefore states the trigger more clearly than “Use for all list retrieval work.” Whether it is selected appropriately still needs to be checked in later tasks.
The conditions for skipping it matter as well. A change limited to retrieval logic has no reason to invoke a schema migration Skill if the schema is untouched. A brief report on a simple fix may not need a procedure for assembling contract and verification records. OpenAI’s product-specific guidance warns that broad or overlapping Skill descriptions can lead to irrelevant instructions being selected.[S1] It should not be read as a guarantee that every model selects Skills in the same way.
Weigh recurring work against upkeep
How often must a task recur before instructions or a Skill are worth creating? The criteria I suggest have no fixed threshold. Four questions are more useful:
- How often do you explain it again? Is it information, such as the execution location, that comes up across tasks, or a condition unique to this request?
- What happens if it is missed? Would you have to repeat a contract review because a relevant test was overlooked, or merely answer one short question?
- Will it remain true for the next task? The location of the authoritative response contract may stay stable; an exception for this request may disappear immediately.
- What will setup and upkeep cost? Count the time spent writing instructions, diagnosing why they were applied incorrectly, and updating a link when the authoritative source moves.
These are my criteria for making the investment decision, not an official product formula. Token counts alone hide other costs. Consider how often you must explain the same thing, when you have to intervene, how much rework a wrong result causes, and the effort and quality involved in reviewing the final result. OpenAI’s advice to revisit accumulated instructions and evaluate both Skill invocation and outcomes also gives a reason to inspect what happens after a setup is created.[S1][S2]
For an occasional list retrieval change, putting its conditions in the request prompt may be enough. If the response contract differs by caller, turning those differences into a permanent general rule could make exceptions harder to manage. If a short prompt already communicates the scope clearly and lets you review the result, keeping it is a sound investment decision.
On the other hand, if you repeatedly have to point to the same authoritative source and redo contract reviews whenever it is missed, one small link in the project instructions may earn its place. If organizing review records genuinely follows a consistent order across several tasks, a Skill may then be worth the effort. In both cases, the benefit is a prediction, not an observed improvement.
Change one thing, then check whether it earns its place
To assess a setup change, first define the problem you will compare it against. For example, if the problem is “I keep having to explain where the relevant tests are,” check on a similar later task whether those tests were found and how much effort reviewing the result took. For a Skill candidate, check both whether it was applied when needed and whether it was invoked for tasks that did not need it. OpenAI’s article on evaluating Skills recommends setting criteria in advance for expected invocation, missed steps, unnecessary execution, and outputs.[S2]
Then change one thing. If you add an instruction pointing to the authoritative location of the tests, do not create a review Skill at the same time. On similar real tasks afterward, look for whether the needed information is found, unrelated procedures intrude, or important conditions are missed. Contracts that code can judge, such as sorting and pagination behavior, still need relevant test or script inputs, assertions, and actual execution results.
If the change shows no benefit, or interpreting and fixing the setup costs more than it saves, narrow or remove it. If a Skill keeps appearing on tasks that do not need it, tighten its name and description or reconsider whether the work is a repeatable procedure at all. When the model or project changes, revisit why each surviving rule was created, whether its authoritative source is still correct, and whether it duplicates another instruction. Official advice to review accumulated instructions is a useful starting point for that check.[S1] Tidying up old workflow guidance is distinct from removing permission, approval, or security boundaries.
The lasting skill is knowing where information belongs and how to check it
As tools change, instructions and Skills that fit today may need another look. What remains useful is the ability to spot recurring work, choose where its information belongs, separate contracts that code can check from judgments a person must review, and reassess the upkeep. This is a case for investing when the burden actually recurs, rather than a conclusion that everyone must master every feature.
Think of one thing you have explained several times recently. Is it needed only for this request, stable across the project, a defined repeatable procedure, or an executable contract? The answer may justify one small change—or keeping your current prompt. The next article, “What Should I Learn Now, and What Can I Put Off?”, will carry this judgment into learning priorities.
Sources
- [S1] Rethinking skills and prompts for GPT-6 Astra | Eric Provencher, OpenAI Developers | Published 2026-09-11; accessed 2026-10-05 | https://developers.openai.com/blog/rethinking-skills-and-prompts-for-gpt-6-astra ↩
- [S2] Testing Agent Skills Systematically with Evals | Dominik Kundel and Gabriel Chua, OpenAI Developers | Published 2026-01-22; accessed 2026-10-05 | https://developers.openai.com/blog/eval-skills ↩
- [S3] Iterating Development Workflows with Codex | Tony Alaniz, OpenAI Developers | Published 2026-08-03; accessed 2026-10-05 | https://developers.openai.com/cookbook/examples/codex/iterating-development-workflows-with-codex ↩
Report an error or share feedback
Open a draft with this article’s title and URL. Review the message and recipient before sending.
Open email draftIf no email app opens, copy these details into your usual email service.
Related posts
AI Getting the Most from AI Coding Agents: Design the Delegation, Not Just the Tool Comparison
Compare AI coding agents and their controls, then plan task boundaries, verification, approvals, and recovery for practical delegation.
Today I Learned 2026-10-06 Practical Claude Code Commands for Developers
The habits heavy Claude Code users keep recommending, and the commands and shortcuts that put them into practice, from the basics to recent additions, checked against the official docs.