The plumbing beat the model
By Dino Nokic · SEP 2026 · 4 MIN READA study published on 3 September went out to 285 organizations and asked a question almost nobody asks out loud: not which model are you running, but what have you built around it.
49% of the organizations with mature context engineering also qualified as AI leaders. Among everybody else, 12% did. Four to one, on the strength of the wiring alone.
That is the whole finding, and it is worth sitting with, because it is the opposite of how the money gets spent.
What actually got measured
The interesting part is what put an organization in the leading group. Six things. Data integration. Workflow orchestration. Retrieval methods. Federated metadata. Prompt engineering. A semantic layer.
One of those six is about the model. The other five are plumbing — the unglamorous business of getting the right information to the right place, in a form something can use, with a record of where it came from.
Then the obstacles, ranked by the same 285. Data quality and preparation, 49%. Model limitations, 29%. Governance gaps, 25%. Data freshness, 23%.
Read that ranking twice. The model is the third-largest problem in the room. It is beaten by the state of the data and it is nearly matched by not knowing who is allowed to do what. Everyone I talk to is shopping in the 29%.
Not exotic. Just unfinished.
The number that stops me is a different one. 42% of respondents already qualified as context leaders.
This is not a rare capability held by four laboratories on the west coast. Nearly half of an ordinary sample had done the work. Which means the four-to-one gap is not a gap between the funded and the unfunded — it is a gap between the organizations that finished six boring jobs and the organizations that started them and moved on to something more interesting.
I wrote here recently about whether the single number in your record is true. This is the layer above that. You can have honest fields and still lose, if nothing can find them, if two systems disagree about what the word means, and if nobody can say where the figure came from.
What these six are, on a real floor
Strip the vocabulary off and every one of them is a question you can ask about a transportation and logistics operation this afternoon.
Retrieval is: when somebody needs the rate confirmation from four weeks ago, how long does that take, and does it come back the right one. Semantic layer is: does delivered mean the same thing in the tracking system that it means in the billing system. It usually does not. One means wheels stopped, the other means paperwork cleared, and the two are a day apart.
Federated metadata is: can anybody tell you where a number was born — measured by a device, or typed by a person at the end of a shift. Orchestration is: when the answer arrives, does anything actually happen, or does it land in an inbox behind forty other things.
None of that is a model problem. All of it is a seam problem. Seams are where the work goes to die, and they are invisible on an org chart because no single person owns both sides.
Where to start on Monday
Grade yourself on the six, honestly. Write them down the left of a page and mark each one: nothing, started, actually finished. Do it with the people who touch the work, not the people who describe it. The half-finished ones are your answer, and this costs you a morning.
Do the semantic layer first. Define the words before you buy anything that reads them. Take the ten terms your decisions rest on — delivered, available, on time, closed — and force one definition per term across every system. It is the cheapest of the six and it unlocks the rest, because until it is done, everything downstream is comparing two different things and calling it analysis.
Then close the seam that gets argued about twice a week. Not the biggest one. The one that already generates a recurring argument, because that argument is a free measurement you have been ignoring — it tells you exactly where two sides of a handoff disagree.
Keep the human gate on while you do it. Governance gaps sat at 25% in that same ranking, ahead of freshness. A gate is not hesitation. It is the thing that lets you find out you were wrong while it is still cheap.
When we have run this order of work on an operation, the returns showed up in the seams and not in the tooling. Inbound reaction went from three days and more down to about sixty seconds. Daily escalations came down from seven to ten a day to three to five inside six weeks. Roughly 20% more efficiency overall — from six boring jobs, done properly, in the right order.
Everyone is shopping in the 29%. The four-to-one gap is sitting in the other five things, and none of them are for sale.
Sources: BARC, “Context Engineering for Agentic AI: Architecture, Use Cases, and Principles for Success” — survey of 285 organizations (3 Sep 2026) · 01net report on the BARC study findings (3 Sep 2026)