Essay
v1.0
When Systems Upgrade Systems
Systems that identify their own deficiencies — and the governance they demand.
A note from the author
This is probably the idea I am most excited about in the current Revenue Labs architecture. We built a simple feedback mechanism where someone can disagree with an agent. That led us to a bigger question: instead of only correcting the output, could the system diagnose why it was wrong? Missing State? Weak Memory? A bad Skill? A missing tool? An incomplete ontology? Follow that far enough and the product starts participating in improving the system that produced the product. I do not mean uncontrolled recursive AI. I mean governed systems that can observe their own deficiencies, propose improvements and learn through operation. This essay explores what that could become.
Most organisations improve through deliberate human intervention. A problem appears. Someone notices. A manager investigates. RevOps changes a process. A new playbook is written. Training is updated. A workflow is rebuilt. Eventually, the organisation operates differently. This is how companies learn.
It is also slow, uneven and difficult to connect back to outcomes. The process exists across people, documents, software and habit. The organisation can usually see that something went wrong. It is much harder to identify exactly which part of the operating model produced the failure.
AI-native systems create another possibility. If the operating system maintains State, applies Memory, executes Skills, preserves Decision Traces and observes outcomes, it can begin to inspect the relationship between how work was performed and what happened afterwards.
That changes the question. Not only: Can the system perform this capability? But: Can the system identify why the capability is weak, and improve the conditions under which it operates?
That is the beginning of a self-upgrading organisation.
Work creates evidence about the system
Consider a Deal Assessment capability. It evaluates an opportunity. The system believes the deal is healthy. A manager agrees. The opportunity later slips.
One failed prediction tells us relatively little. Commercial outcomes are noisy. But imagine the same pattern occurs repeatedly. Deals with a particular shape are consistently rated too highly.
Now the operating system has a problem worth investigating. There are several possible causes. The Live State may be incomplete. A meaningful signal may not be represented. The company’s Memory may contain a weak definition of qualification. The Deal Assessment Skill may apply the right knowledge badly. The agent may be using the wrong tool. The ontology may not represent something experienced managers understand intuitively. Or the organisation itself may hold an incorrect belief about what healthy deals look like.
These are very different failures. A traditional system sees the outcome. An AI-native system can increasingly inspect the architecture that produced it.
This is important because every execution creates evidence about more than the task. It creates evidence about the quality of the operating model.
The first upgrade loop is human
The simplest version already exists. A system produces an output. A human gives feedback. Thumbs up. Thumbs down. Accept. Reject. Edit.
This is useful. But the signal is shallow. A rejection tells us that the user disagreed. It does not tell us why.
An AI-native learning loop should go further. The system can ask: Was relevant State missing? Did the user possess context the system did not? Was the Memory wrong? Did the Skill fail to apply the Memory correctly? Was the recommendation sound but poorly explained? Was the action mistimed? Was the human simply wrong?
The feedback becomes diagnostic. The human is no longer only approving an output. They are helping identify which layer of the system needs improvement. That is a much more valuable role.
Overrides are evidence about architecture
Human overrides become particularly important. Suppose experienced managers repeatedly override a Deal Risk capability. The naive response is to improve the model. That may be the wrong fix.
Perhaps every override involves a specific kind of stakeholder relationship the ontology does not represent. Perhaps the managers rely on evidence from private executive conversations that never enters the system. Perhaps the company’s definition of champion strength is too simple. Perhaps the Skill is ignoring a signal already present in State.
The disagreement is not merely evidence that the AI made a bad decision. It is evidence that the operating model may be incomplete.
Repeated overrides create a map of where organisational judgement remains outside the system. That makes disagreement valuable.
The system can increasingly ask not only: “Was I wrong?” but: “What would I have needed to understand in order to make the better decision?”
Missing State is one class of failure
Some problems are failures of representation. The system cannot reason about something it cannot see. A deal repeatedly slips because procurement involvement matters, but procurement is represented only as another contact. Customer risk is missed because stakeholder change is not connected to account health. Prospecting underperforms because the system records a leadership change but not whether the new executive has a history of buying similar products.
In each case, the problem is not the Skill. The relevant concept is absent or weakly represented in State.
The system can identify this pattern. Which missing variables appear repeatedly in human explanations? Which external evidence do people manually retrieve before overriding the system? Which decisions have low confidence because the same context is consistently absent?
This can produce an upgrade recommendation: The commercial model needs a richer representation of procurement involvement. Or: Stakeholder mobility should become part of customer State.
The system begins participating in the evolution of its own representation of reality.
Memory can be wrong
Other failures come from judgement. The system has the relevant facts. It interprets them badly because the organisation’s Memory is incomplete or outdated.
Perhaps the company believes that firms above a particular size are its strongest ICP. Outcome data increasingly suggests that organisational structure matters more than headcount. Perhaps the sales methodology treats Economic Buyer identification as sufficient evidence, while successful deals show that meaningful engagement is the stronger predictor. Perhaps a positioning assumption that worked twelve months ago no longer resonates.
The system can compare: what the organisation says it believes, how capabilities apply those beliefs, and what happens afterwards.
This makes organisational Memory testable. Not every belief should be mechanically optimised against short-term outcomes. Strategy contains deliberate choices. Markets change. Causality is difficult.
But the system can identify tension. “The current ICP Memory predicts these accounts should outperform. Over the last six months, they have not.”
That is valuable. The system is not making strategy. It is showing the organisation where reality is challenging its strategy.
Skills can be debugged
Sometimes the knowledge is right and the method is wrong. The company has a strong definition of a champion. The Deal Assessment Skill fails to apply it consistently.
The Account Research Skill uses too much generic market information and too little company-specific evidence. The Meeting Preparation Skill produces thorough briefings but consistently misses unresolved customer commitments.
These are Skill-level problems. Because Skills are executable and versioned, they can be evaluated more like software. Which steps contribute to useful outcomes? Where do errors concentrate? Which conditions produce failure? Which human edits recur? Does one version outperform another?
The system can begin proposing modifications. Add this evidence source. Change this sequence. Increase the weight of this condition. Introduce an escalation when uncertainty exceeds a threshold. Remove a step that adds latency without improving outcomes.
This is process improvement becoming more empirical.
Tools can be the bottleneck
A capability can also fail because it cannot act or observe effectively. The system knows that product usage would improve Customer Risk assessment. It has no access to product data.
The agent needs to verify a commercial commitment in email. The connector does not exist. A Skill repeatedly asks for information that no available tool can retrieve.
Traditional software failures of this kind often surface through user frustration. An AI-native operating system can identify them systematically.
“This capability repeatedly reaches low-confidence conclusions because support data is unavailable.” Or: “Managers manually consult billing before overriding Renewal Risk recommendations.”
The system has identified a missing tool. Now integration prioritisation can be based on observed capability gaps rather than generic feature requests. That changes the product-development loop.
Ontology can be the hidden problem
Some failures sit even deeper. The system may have data, Memory, Skills and tools but still lack the conceptual structure required to reason well.
Suppose experienced salespeople distinguish between: a supportive contact, a mobiliser, a politically powerful sponsor, and an Economic Buyer.
The system represents all four as “stakeholder”. No amount of better prompting fully solves the problem. The ontology is too weak.
The operating system has not represented an important distinction in the commercial world. This is one of the harder upgrade classes because ontology determines what the system is capable of seeing.
It also explains why repeated failures can be useful. If the same unrepresented concept appears across many human corrections, the system can propose that the commercial model itself needs to evolve.
This begins to resemble scientific model-building. The system observes phenomena its current representation does not explain well. It proposes a richer model.
Repeated manual work can reveal missing capabilities
The system can also learn from what people repeatedly do outside it. Imagine managers continually perform the same analysis before strategic deal reviews.
They gather several pieces of State. Apply similar judgement. Produce similar recommendations. No formal capability exists. The behaviour repeats.
This is evidence. The system can recognise the pattern and ask: Should this recurring work become an explicit capability?
That is a powerful upgrade mechanism. Today, companies identify automation opportunities through interviews, process mapping and operations analysis.
An AI-native operating layer can observe repeated patterns directly. What work do humans repeatedly perform? Which tools do they use? Which information do they gather? What decisions follow? How consistent is the method?
The system can propose candidate capabilities. The company starts discovering its own automation roadmap through operation.
A system can build a system
This is where the idea becomes more interesting. Suppose a user repeatedly asks Revenue Labs to perform a particular type of analysis. The system does not simply answer. It notices that the request is recurring.
It identifies the State required. The relevant Memory. The method being used. The tools invoked. The output format. The human corrections.
Eventually it can propose: “This appears to be a recurring organisational capability. Would you like to formalise it?”
The resulting Skill can be generated from previous successful executions. The capability can be tested. Governance can be added. It becomes reusable.
Now the system has helped turn ad hoc work into infrastructure. An agent has effectively helped build another part of the agentic system.
This is what systems upgrading systems begins to mean in practice. Not an uncontrolled recursive AI rewriting itself. A governed operating system observing repeated work and helping the organisation make that work explicit, reusable and improvable.
The upgrade surface becomes a product
This changes the user experience. Today, software primarily surfaces work. Tasks. Insights. Reports. Alerts.
An upgrading system also surfaces proposals about the system itself. For example: “Account Prioritisation is frequently overridden for Series C infrastructure companies. These accounts convert 2.4× above the current model’s expectation. Consider updating ICP Memory.”
Or: “Deal Assessment has low confidence whenever procurement appears before Economic Buyer engagement. Managers consistently treat this pattern as high risk. Add procurement sequencing to the Skill?”
Or: “Pipeline Review manually invokes the same stakeholder analysis in 73% of strategic deals. Consider adding Stakeholder Risk as a reusable Skill.”
These are not ordinary recommendations. They are recommendations about the operating model. That is a new class of software experience.
The system begins exposing where it thinks the company itself can improve.
Upgrades need evidence
This creates obvious governance requirements. The system should not say: “I think we should change the ICP.” It should show the evidence.
Which decisions revealed the problem? Which traces were inspected? Which outcomes support the conclusion? How large is the sample? Which alternative explanations exist? Which capabilities would be affected? What is the expected impact?
This is where Decision Traces become essential. They provide lineage. The proposed upgrade can point backwards through the operating history that motivated it.
This makes system improvement inspectable. And it allows humans to reason about the recommendation rather than simply trusting the model.
Upgrades should be proportional to consequence
Not all changes deserve the same governance. A Skill changes the ordering of information in a meeting briefing. Low consequence.
A system proposes a new customer-risk signal. Moderate consequence. A change to pricing policy. High consequence. A change to ICP that redirects millions of pounds of commercial attention. Very high consequence.
The operating system should reflect this. Some upgrades can be automatic. Some can run as experiments. Some require RevOps approval. Some require functional leadership. Some require executive governance.
This is consistent with the broader principle of bounded autonomy. The system can increasingly participate in its own improvement without becoming self-governing. That distinction is central.
Upgrades can be tested before they propagate
Versioned Memory and Skills make experimentation possible. Suppose the system proposes changing Account Prioritisation. Instead of updating the entire organisation immediately, the new version can run in parallel.
Compare recommendations. Observe outcomes. Evaluate human acceptance. Test a segment. Then decide whether to promote the change.
This creates something closer to deployment discipline for the operating model. Version. Evaluate. Canary. Approve. Roll out. Rollback if necessary.
Commercial process starts inheriting useful ideas from software engineering. Not because organisations become code. Because once operating logic becomes executable, changes to that logic deserve similar care.
The organisation can run competing hypotheses
This creates another possibility. Companies often disagree about strategy. Which segment is stronger? Which signals matter? What constitutes a qualified opportunity?
Today these disagreements are often resolved through seniority, debate or retrospective analysis. An executable operating model can make some of them testable.
Memory A represents one ICP hypothesis. Memory B represents another. Capabilities apply each under controlled conditions. Decision Traces preserve the reasoning. Outcomes accumulate.
The organisation learns. Again, not every strategic question can or should be reduced to an experiment. Markets are reflexive. Samples are imperfect. Long-term strategy often requires acting before evidence is conclusive.
But the boundary between belief and testable operating hypothesis can move. That makes the company more empirical.
The system learns at different speeds
This is important. Not every layer should update at the same cadence. State may change continuously. A low-risk Skill might improve weekly. A qualification methodology might change quarterly. ICP may change more deliberately. Core strategy might remain stable for much longer.
A self-upgrading organisation therefore needs learning rates by layer. The system should know which things are fluid and which require stability. Otherwise rapid optimisation can create organisational thrashing.
This is a general problem in adaptive systems. A system that reacts to every new piece of evidence becomes unstable. Learning therefore requires memory, but also damping.
Enough responsiveness to adapt. Enough stability to preserve coherent direction. That trade-off will become an important part of AI-native governance.
Local optimisation can damage the whole system
There is another danger. A capability can improve locally while making the organisation worse globally. A Prospecting Skill optimises for reply rate. It learns to target easier-to-reach accounts. Pipeline quality falls.
A Meeting Follow-up capability optimises for rapid next steps. Customers feel pressured. A Renewal capability optimises retention through excessive discounting. Short-term outcomes improve. Economic value falls.
This is why self-upgrading systems need objectives above the capability level. The organisation is not a collection of independent optimisers. Capabilities share resources and interact. Improving the whole system requires understanding those relationships.
The prospecting system must understand pipeline quality. The sales system must understand customer outcomes. The customer system must understand lifetime value.
The operating model needs hierarchy of objectives. That is another reason leadership remains central. AI can optimise methods. Humans still decide what the company is trying to become.
The system needs a model of objectives
We have spent a lot of time discussing State, Memory and Skills. A self-upgrading system makes objectives more explicit too. What is this capability trying to improve? Meeting conversion? Pipeline quality? Forecast accuracy? Retention? Long-term customer value? Learning in a new market?
Without clear objectives, improvement has no direction. And objectives often conflict. Short-term revenue versus long-term positioning. Automation versus customer experience. Efficiency versus learning. Consistency versus experimentation.
The system needs enough representation of those objectives and constraints to understand whether an upgrade is actually an improvement. This pushes AI-native architecture closer to the operating model of the company itself.
Improvement becomes observable
Traditional process improvement is often difficult to measure. A new playbook is launched. Six months later, performance improved. Did the playbook cause it?
Perhaps. Or the team changed. The market improved. The product got better. A competitor weakened.
Executable capabilities create a more detailed record. Which version ran? For which situations? What changed? How did users respond? What happened afterwards?
This does not eliminate causal uncertainty. But it improves the evidence available. The organisation can begin accumulating a history of operating-model changes and their consequences. That history itself becomes valuable.
The company learns not only from commercial decisions. It learns from changes to how it makes commercial decisions. That is a higher-order learning loop.
First-order and second-order learning
This suggests a useful distinction. First-order learning improves a decision within the current system. The agent gets better at assessing deals.
Second-order learning improves the system that produces those decisions. The organisation changes the Deal Assessment Skill. Or the definition of a champion. Or the ontology. Or the data available.
This is where system upgrading becomes genuinely interesting. The organisation is not only learning what to do. It is learning how it should learn and decide.
In systems theory, this is a deeper level of adaptation. In organisational terms, it is the difference between improving performance and improving the operating model.
Human judgement moves upward again
This changes human roles. At first, humans perform the work. Then AI performs more of the work and humans supervise. As the system becomes more reliable, humans increasingly govern the system’s improvement.
Should this Memory change? Is this apparent correlation strategically meaningful? Should this Skill upgrade propagate? Is the system optimising the wrong outcome? Is this exception evidence of a deeper pattern or merely an exception?
These are higher-order judgement questions. Humans move from: doing the task, to reviewing the task, to improving the capability, to governing how the system improves capabilities.
This is another form of moving up the stack.
RevOps becomes the governor of upgrades
This is where the future RevOps role becomes particularly clear. RevOps does not need to manually discover every system deficiency. The operating layer can surface them.
RevOps evaluates the recommendation. Which layer is affected? How strong is the evidence? Who owns the underlying judgement? What capabilities will inherit the change? What experiment should run? What governance is required?
RevOps increasingly manages the release process of the commercial operating model. That is a useful analogy. Memory has versions. Skills have versions. Capabilities have evaluations. Upgrades have evidence. Changes propagate. Some can be rolled back.
Commercial operations starts to look much more like a continuously evolving system.
External knowledge can upgrade the system too
There is another direction for improvement. The system does not need to learn only from internal outcomes. External knowledge can enter the loop.
A new industry framework. Research. A sales methodology. Partner expertise. Benchmark data. Updated market knowledge. A specialist Skill.
The operating layer can compare external knowledge against internal context. Does this methodology apply here? Does it contradict current Memory? Could it improve an existing Skill?
This creates a richer model of organisational learning. Internal experience compounds. External knowledge extends it.
The company does not simply consume content. It can increasingly integrate useful external knowledge into executable operating logic. This may eventually create markets for high-quality Skills and Memory modules. Expertise becomes infrastructure.
Upgrading creates a different moat
This has strategic consequences. Imagine two companies using similar foundation models. Both start with comparable capabilities. One treats the AI system as static software. The other has a robust upgrade loop.
Its human corrections become traces. Its outcomes become evaluations. Its Skills improve. Its Memory evolves. Missing capabilities are discovered. External expertise is integrated.
After a year, the two systems are no longer comparable. After five years, the difference can become substantial.
The advantage is not that one company bought a better model. Both can access better models. The advantage is the rate at which each organisation turns experience into improved capability.
That is learning rate again. And learning rate compounds.
The product starts designing the product
There is a strange implication here. Most software is designed by product teams. Users provide feedback. The team decides what to build.
An AI-native operating system can participate more directly in that loop. It can see: which capabilities are requested repeatedly, where existing capabilities fail, which tools users manually add, which Memory is missing, which Skills users repeatedly modify, where attention concentrates, which outcomes remain poor.
The product can produce a much richer representation of its own deficiencies. Eventually, it can propose solutions.
New Skill. New tool integration. New ontology concept. New capability. The product team moves from guessing at some needs to governing a system that can increasingly articulate its own gaps.
This does not eliminate Product. It gives Product a new collaborator. The product starts helping design the product.
The limit is not intelligence alone
There is a tendency to imagine that more capable models automatically lead to self-improving organisations. They do not. The architecture matters.
Without State, the system does not know enough about the environment. Without Memory, it lacks stable organisational judgement. Without Skills, methods are difficult to isolate and improve. Without Decision Traces, causality disappears. Without outcome data, there is no feedback. Without versioning, changes cannot be evaluated. Without governance, improvement becomes dangerous. Without objectives, optimisation becomes directionless.
Self-upgrading is therefore not a model feature. It is a system property. That distinction matters enormously.
From automation to adaptation
We can now see a broader progression. Traditional software automates. Given a process, execute it consistently.
AI-assisted software accelerates. Given work, perform more of it. AI-native systems adapt. Observe State. Apply judgement. Choose capabilities. Act. Learn from outcomes.
Self-upgrading systems go one step further. They improve the machinery of adaptation itself. The distinction is:
Automation changes what happens. Adaptation changes what should happen. Self-upgrading changes how the system decides what should happen.
That is a much deeper transformation.
The company becomes recursive
At this point, the operating loop becomes recursive. The company operates through its capabilities. Those capabilities generate outcomes.
The outcomes reveal something about the quality of the operating model. The operating model changes. The changed operating model produces different future outcomes.
The company is now acting on the world and on itself. That is what makes this stage structurally different.
The organisation is not merely a system with feedback. It is a system capable of modifying the mechanisms through which it responds to feedback. This is where the language of the learning organisation becomes literal rather than metaphorical.
The frontier is governed self-improvement
There is a tempting science-fiction version of this future. Autonomous agents rewriting themselves continuously. Recursive systems operating without people. Companies running themselves.
I think that framing obscures the much more interesting near-term transition. The frontier is not uncontrolled self-modification. It is governed self-improvement.
A system capable of observing its own performance. Identifying deficiencies. Diagnosing which layer is responsible. Proposing improvements. Testing them. Showing the evidence. Escalating consequential changes. Propagating approved upgrades. Learning from what happens next.
That is already a profound change in how companies can operate. It means organisational learning no longer needs to live entirely outside the software as an occasional management activity. Learning becomes part of the architecture.
When systems upgrade systems
The original promise of software was consistency. Encode a process and machines perform it the same way every time. The promise of AI is intelligence. Machines can increasingly interpret changing situations and perform work that previously required human judgement.
The next step is improvement. The system begins recognising that the way it currently understands or performs the work is not good enough. It identifies why. And it participates in changing itself.
That creates a new operating loop: State → Judgement → Capability → Action → Trace → Outcome → Diagnosis → Upgrade.
Then the loop begins again. With better State. Better Memory. Better Skills. Better capabilities.
The organisation improves because it operates. At that point, software is no longer simply supporting the operating model. It is participating in its evolution.
That is when systems begin upgrading systems.
This essay is versioned. Where our thinking develops materially, we will update the version and explain why - the revision history is preserved, not polished away.

