OpenAI and Anthropic recently published migration guidance for their newest models, and both carry an unusually candid message for developers: the prompts and skills that earned good results from the last model can hold back the new one.
Anthropic’s guide for Claude Fable 5 notes that skills written for prior models “are often too prescriptive for Claude Fable 5 and can degrade output quality.” OpenAI’s guide for GPT-6 Astra strongly recommends auditing skills and files like AGENTS.md, because the new model responds more closely to the instructions it finds there. Across both model families, the conclusion is the same, and instructions tuned for one model can become the bottleneck for the next.
Every line you give an agent exists for one of two reasons, either because of your domain or because of the model. Domain content covers facts about your systems, boundaries set by business, and the intent behind a task, and it stays true when the model changes.
Model content is a workaround for a weakness, so it expires when the weakness does. Sometimes it even works against you: GPT-6 Astra asks more clarifying questions than earlier OpenAI models, so an instruction that once encouraged a model to check in can now stall one that already does.
The difference is worth designing around. Instructions depreciate and context appreciates. A workaround loses value once the weakness it patched disappears, while a fact gains value with every release because a stronger model extracts more from it.
Most current advice focuses on prompt hygiene, which means auditing files and deleting workarounds. The larger opportunity is deciding where your domain knowledge should live, so agents improve with every model release without anyone rewriting instructions. For media, that answer is concrete, and it matters even more because the prompt was never the right home for media knowledge.
An agent can open your code, configs, and docs and reason about them directly, because in text the artifact and the knowledge are the same thing.
Media works differently. A vision model can look at an image and describe what’s in it, but the questions that matter in production live outside the pixels. Who holds the rights? Which campaign does it belong to? Is it an approved master or a stale duplicate? Where does it already live?
That knowledge often sits in a DAM that agents don’t query, a spreadsheet in a shared folder, or in the memory of a designer who has since moved on. Wherever it lives, the agent can’t reach it.
Teams compensate the only way they can: in the prompt. “Hero images are 1,200 pixels wide.” “Convert to WebP.” “Check usage rights before publishing.” But if you run the “domain-or-model” test on those lines, you can be sure something odd shows up.
The facts are domain content, but they’re delivered like model content, maintained by hand, trapped in one prompt, invisible to every other agent, and stale the moment anything changes. The durable fix is media that carries its own context.
For an agent to reason about media before acting on it, the asset needs to answer four kinds of questions on its own.
- Identity. What is this asset? Technical facts like format and dimensions sit alongside semantic ones, including AI-generated tags, captions, and alt text. Structured metadata fields tie the asset to your business, such as the product SKU, campaign, and market.Provenance recorded at upload shows where the asset came from, including which model generated it. With that information, an agent picks the right image based on what the asset says about itself.
Provenance recorded at upload shows where the asset came from, including which model generated it. With that information, an agent picks the right image based on what the asset says about itself.
- Relationships. What is this asset connected to? The source master, the variants derived from it, and the pages and channels each one serves are all on record. When an agent updates the source, it already knows the dependents, so changes land cleanly across every template.
- Governance. What’s allowed? Usage rights, license windows, moderation status, brand rules, and approved contexts travel with the asset as guardrails. Every agent inherits them automatically so that no prompt has to remember them.
- Presentation. How should this asset show up? The agent states its intent, such as hero slow, an email header, or a product thumbnail, and transformations resolve the right crop, format, and quality for that destination. This replaces the “resize to 1200 pixels, convert to WebP” prompts, which is exactly the kind of instruction that ages with each model generation. Intent holds its value from one model to the next.
Once assets carry their own context, an agent’s capabilities come from the media layer instead of whatever the prompt remembered to include.
- Find media by meaning and permission. The agent searches for approved lifestyle shots with cleared rights for a specific market. When the brief is a reference image, visual search finds assets that look like it.
- Ship the right rendition anywhere. The agent states the destination, and the correct size, format, and quality follow, including destinations the prompt never anticipated.
- Check before publishing. The agent reads the rights window, moderation status, and brand rules, and flags a conflict for a human to review.
- Update once and propagate everywhere. Replace the source asset, and relationships identify every variant and place it lives.
- Turn generated files into production assets. A model returns a file with no context and limited lifespan. Uploaded as an asset, it starts building identity, relationships, and governance from day one.
Each of these is an ordinary production task. In a context-rich setup, each one keeps working as models change, with no one maintaining a prompt to support it.
Cloudinary treats every asset as a structured media object. Auto-tagging and AI captioning fill in identity, structured metadata ties assets to your business, moderation and access rules carry governance, and transformations handle presentation from a statement of intent.
Most of that context takes little manual effort. Enrichment runs at ingestion: upload an asset and auto-tagging and AI captioning describe it on their own. Moderation runs automatically on upload with a single parameter. Access rules travel with the asset instead of living in application code, and structured metadata turns business context into fields the agent can query. The cost of adding context drops to the cost of an upload.
Agents connect through official Cloudinary MCP servers and get queryable context back. We apply the same domain-or-model test to our own agent tooling, keeping facts, boundaries, and intent in the files our agents read and pruning model workarounds with every release.
Media built this way keeps improving because of where models are headed. With each release, a smarter model pulls more from the same object. It catches the rights conflict that the last model skipped, picks a better rendition for a context you never specified, and connects the campaign metadata to a task you didn’t anticipate. A context-rich media setup collects each upgrade automatically, while a prompt-driven one needs a rewrite every generation.
For years, the craft was writing better instructions for the model you had. With media, the craft is building assets that explain themselves — to whichever model shows up next.









