For three years the standard enterprise AI pattern was retrieval-augmented generation. You chunk your documents, embed them, retrieve the most relevant pieces, stuff them into a 100,000-token context window and generate an answer. We built whole stacks around it.

Anthropic's Claude 4 family, with its one-million-token context window now widely available, has made parts of that stack optional. Not all of it, but more than I expected when I started testing.

What changed

A million tokens is roughly 750,000 words. A typical enterprise contract is about 10,000 words, and a full year of board minutes might be 100,000. A medium-sized codebase fits comfortably. So does a complete employee handbook with its policies, procedures and FAQs.

Size is only half of it. The model can now hold all of that at once and reason across it without losing detail. Recall on million-token contexts in Claude 4 is good, and the "lost in the middle" problem that plagued earlier long-context models has largely gone.

Add prompt caching at the provider level and sending the same large context across many queries goes from too expensive to routine.

Where long context wins outright

Document analysis

For tasks like "answer questions about this contract" or "summarise this 200-page report", long context now beats RAG. No chunking step loses detail, and no retrieval step fetches the wrong section. The model reads the whole document and reasons over it directly.

Earlier this year I rebuilt a contract analysis tool that used to chunk agreements into sections and retrieve the relevant clauses. Loading the full contract instead made the answers noticeably more accurate, especially on questions that depended on how different clauses relate to each other.

Understanding a codebase

For coding assistants and review tools, putting a whole repository (or at least the relevant subsystem) into context gives different results from retrieval. The model can trace dependencies, pick up naming conventions and follow patterns across files in a way chunked retrieval struggles with.

Reasoning across a known set of documents

"Compare these three vendor proposals and find the differences in their service-level commitments" is a long-context job. A retrieval approach has to guess which sections to pull from each proposal, and it usually misses something.

Where RAG still wins

Long context doesn't replace retrieval. It's a different tool with a different shape.

Very large corpora

If your knowledge base is truly huge (millions of pages, years of Slack history, everything in SharePoint), you still need retrieval. Long context is great with one complete document. It can't search across everything you own.

High-volume, cost-sensitive work

Sending a big context with every query costs more than retrieving a small relevant slice, even with caching. For customer-facing systems handling millions of queries, the economics still favour RAG.

Data that changes often

Prompt caching breaks when the context changes. If your knowledge updates throughout the day, the caching benefit disappears and long context gets expensive again. Retrieval over a fresh index copes with this naturally.

The hybrid pattern

The teams getting the most from Claude 4's long context aren't choosing between long context and RAG. They're using both.

A light retrieval step narrows millions of documents down to the few hundred pages that might matter. Those pages then go into the context as one complete bundle, and the model reads all of it, reasons across it and answers.

You get the precision of long-context reasoning without loading the whole corpus on every query. It's also sturdier than pure RAG, because retrieval only has to find the right set of documents, not the exact right paragraph.

What it means for architecture

If you're building or rebuilding an enterprise AI system in 2026, the design questions are different. What's the natural unit of context for the task: a document, a project, a customer record? Does it fit in a million tokens? How often does it change, and can you cache it for a session, a day or longer? And where do you really need search-style retrieval, as opposed to places you've been using it out of habit?

For new builds I now start with long context and add RAG where the size of the corpus or the cost analysis calls for it. A year ago I'd have done the opposite.

What I'd revisit

Long context didn't kill RAG. It took RAG out of a lot of places where it was only there by default. If you built your AI stack in 2023 or 2024, some of those decisions were right at the time and aren't any more. The chunking strategy is usually the first thing worth another look.

Wondering whether long context could simplify your current setup? Send me a note and I'll give you a second opinion.