Rollingstone Revelations - Koshy's Blog
The weekly series that will make you think, laugh and cry. Don't miss. Bookmark this page
Saturday, August 8, 2026
Sunday, August 2, 2026
Insurance for a Billion: Will AI Make It Fairer, or Just More Precise?
[This article is built around thought I shared in the round table titled “Insuretech for India” at Stride Forward 26]
India built the world's most complete public digital
infrastructure to include people. The next test is whether it uses AI to
protect more of them, or to price the
riskiest ones out.
Insurance is the one financial product designed to work by pooling
strangers together. The healthy subsidise the sick, the lucky subsidise the
unlucky, and everyone buys protection against a future none of them can
predict. That is not a flaw in the model. It is the model. Artificial
intelligence is now very good at predicting exactly who will get sick, who will
crash, and whose house will flood, and that ability, left to run on its own
logic, quietly dismantles the thing that made insurance worth having.
This is the real question hanging over the insurance industry, and it is
sharpest in India, not because India is behind, but because India is unusually
well equipped to take it either way. Over the last decade the country has
assembled the most complete public digital infrastructure in the world: a
billion-scale digital identity, real-time payments, consent-based data sharing,
digital documents and signatures. That stack was built, deliberately, to
include people who markets had left out. The same stack, pointed at insurance
and combined with AI, can be used to include far more people, or to segment them so finely that the ones who
most need cover can no longer afford it. India will have to choose. Most
countries will not get to make that choice as consciously, because they lack
the rails to make either outcome happen at scale.
So it is worth being clear about what is actually at stake, and where the
genuinely hard problem lies — because it is not where most of the industry
conversation puts it.
The easy part: insurance is about to disappear into everything else
Start with the parts that are, by now, close to consensus. The first wave
of insurance technology everywhere, cheaper distribution, faster underwriting,
quicker claims, lower operating cost, is
largely a solved direction of travel. Three shifts follow from it, and they
will define the next several years.
Insurance becomes embedded rather than sold. It stops being an
annual contract you remember to renew and becomes a service that attaches
itself, invisibly, to something else you are already doing. You buy a
two-wheeler and accident cover comes with it. You take a home loan and property
cover is part of the transaction. You book a trip, finance an MSME invoice, buy
farm equipment, or see a doctor on a health platform, and the relevant
protection is simply present. India's digital platforms make this possible at a
scale few markets can match. The right ambition is for insurance to become, in
the phrase I keep coming back to, always present but almost invisible.
AI settles the straightforward claims in near real time. A large
share of claims are simple, honest, and slow only because a human has to look
at them. Those will be assessed and paid in minutes. This matters less as an
efficiency story than as a trust story, which I will come to.
Every citizen carries a portable insurance profile. Identity,
verified financial history, health records shared with consent, property
records, and past claims can travel with the individual rather than being
locked inside one insurer. That lets a person move between insurers without
starting from zero each time, and it lets underwriting happen in real time at a
fraction of today's cost.
None of this is the hard part. It is expensive and fiddly to build, but
the direction is not in doubt and the benefits are real. If this were the whole
story, the correct posture would be enthusiasm and patience.
The barrier that technology alone does not fix
The deeper obstacle in India has never been mainly technological. It is
trust, and it has four distinct faces. People do not reliably know what to buy,
whether a claim will actually be honoured when it matters, whether the premium
they are quoted is fair, or whether the whole process is simple enough to be
worth attempting. Every one of those is a reason someone who should be insured
is not.
Technology helps with each, AI-assisted advice for the first, transparent
pricing for the third, paperless onboarding and cashless claims for the fourth.
But the second, whether claims are honoured, is where technology and trust
actually meet. A claims process that pays honest claims in minutes, visibly and
repeatedly, builds the kind of trust that a marketing campaign cannot. That is
why real-time claims settlement matters more than its efficiency suggests: it
is the mechanism by which an industry with a credibility problem earns credibility
back.
Why India can build rails, not just products
Here is where India's position differs from most markets, and it is worth
stating precisely rather than triumphantly. Elsewhere, the natural unit of
progress is the company: a better insurer, a smarter underwriting model, a
slicker app, each building its own private ecosystem. India has the option to
build the shared layer underneath all of them, common digital rails for insurance in the way
real-time payments became common rails for money, so that insurers compete on
top of shared infrastructure instead of each rebuilding the plumbing.
The reason India can attempt this is that most of the foundation already
exists and is public: verifiable identity, a consent architecture for sharing
financial and health data, digital documents and signatures. Very few countries
have that combination in public hands. The missing layer is insurance itself , the
standards and rails that would let a verified individual be underwritten,
insured, and served across providers with their consent and without friction.
Get that layer right and the cost of issuing and servicing a policy falls far
enough that protecting a low-income family becomes commercially viable rather
than charitable. This is the genuinely globally significant experiment, and it
is why people outside India should be watching it: it is a test of whether
insurance can be run as public infrastructure rather than only as a private
product.
But this is exactly where the capability turns double-edged, because the
same rails that can underwrite a poor family in real time can also price that
family out in real time. The infrastructure is neutral. The choice is not.
The tension at the centre: pooling versus prediction
Traditional insurance rests on risk pooling. AI, fed with rich
personal data, pushes relentlessly toward risk segmentation, pricing each individual according to their own
predicted risk. Taken to their logical ends, these two ideas are in direct
conflict. The better AI becomes at predicting risk, the worse insurance becomes
at sharing it.
Follow the logic to its conclusion. Healthy, low-risk people pay very
little, as they should on pure actuarial grounds. High-risk people face
premiums that climb until cover is effectively out of reach. The result is
quietly perverse: the people who most need protection are the ones priced out
of it. That does not just produce an unfair market. It hollows out the social
purpose of insurance altogether, because a pool that has expelled everyone
likely to claim is no longer performing the one function that justified it.
This is not a hypothetical that regulators have failed to notice. It is
precisely why many jurisdictions already prohibit insurers from using certain
information, genetic test results, disability status, pregnancy, some
pre-existing conditions, various protected characteristics. Those prohibitions
are not technological limits; the data is often perfectly usable. They are
deliberate policy choices to preserve solidarity even when better prediction is
available. AI does not create this dilemma. It sharpens it to a point, by
making near-perfect prediction cheap and universal rather than partial and
expensive.
The better question: predict risk, or prevent it?
There is a way out of the trap, and it comes from asking a different
question. What if AI were used not to charge sick people more, but to make them
less likely to be sick?
Imagine an insurer that continuously observes, with consent, the early
signals, rising blood sugar, worsening
blood pressure, an irregular heart rhythm. The segmentation instinct is to
reprice the moment the risk appears. The alternative is to intervene: a
teleconsultation, nutrition coaching, subsidised medication, a fitness
programme, an early screening. If the intervention works, the person stays
healthier, hospitalisations fall, claims drop, the insurer's costs fall with
them, and society carries a lighter burden of disease. The insurer stops being
a payer of claims and becomes a manager of health outcomes. Prediction is put
to work preventing the loss rather than pricing it.
That reframes the whole debate. The question regulators and builders
should be asking now, at the start of AI adoption and not after the practices
have set, is a simple one with large consequences: should AI be used
primarily to predict risk more accurately, or to reduce risk before it
materialises? An industry that mostly predicts becomes more exclusionary
with every improvement in its models. An industry that mostly prevents becomes
more inclusive as its models improve. Same technology; opposite social result.
The difference is a design choice, and design choices are easiest to make
early.
A question for society, not for the algorithm
Insurance has always balanced two principles that pull against each
other: actuarial fairness, which says each person should pay according to their
own risk, and social solidarity, which says a pool should absorb risk on behalf
of those who draw the unlucky number. For most of the industry's history the
balance was set by ignorance, insurers
simply could not price individuals finely enough to fully abandon the pool. AI
removes that ignorance. It will make actuarial fairness very nearly perfect.
Which means the balance can no longer be left to accident. The real
question is no longer what insurers are able to price, but how much
solidarity a society chooses to keep once perfect pricing is possible. That is
not a question an algorithm can answer. It is a question for society, and it
has to be answered on purpose.
India is unusually well placed to lead that conversation, and not by
coincidence. Its digital public infrastructure gives it prediction capabilities
that few countries can match, and its public policy has, through that same
infrastructure, consistently chosen to expand inclusion rather than optimise
markets for their own sake. That combination is rare: the technical capacity to
segment perfectly, paired with a stated preference for including everyone. The
test for Indian insurance technology — and the reason the rest of the world has
a stake in how it goes — is whether it uses AI to expand protection or
merely to refine pricing.
India's last financial inclusion story was about opening bank accounts,
and it largely succeeded. The next one will not be about accounts at all. It
will be about how many lives we choose to protect once we finally have the
tools to protect, or to exclude, every one of them.
The better a machine gets at predicting who will suffer, the more deliberately a society must decide to stand with them anyway.
Friday, July 24, 2026
The Queen or the Swarm: Why AI’s Future Depends on Who Gets to Learn
The detailed version of my oped in Transcontinental Times
Every species that has ever competed for resources on this planet has
done so with roughly the same toolkit: strength, speed, camouflage, venom,
numbers. Humans are unremarkable on most of these axes. We are slower than a
cheetah, weaker than a chimpanzee, blinder in the dark than an owl, and worse
at smelling danger than almost anything with a snout. What we do better than
any other species, by an order of magnitude no other animal comes close to, is
cooperate with strangers at scale.
A wolf pack cooperates. So does a beehive. But a wolf will not lay down
its life for a wolf it has never met from a pack three mountains away,
coordinated by a shared story none of them witnessed. Humans do this routinely.
We show up at war memorials for people we never knew, buy shares in companies
run by executives we’ve never met, and hand savings to banks based on nothing
but a shared belief that the institution will honor its promises. Yuval
Harari’s core insight in Sapiens, that humans are unique in our capacity to
organize around shared fictions: nations, currencies, corporations, human
rights, describes exactly this. None of these things exist as physical objects.
They exist because enough people agree to act as if they do, and that agreement
is what lets seven strangers organize a supply chain across four continents.
That capacity has a name in organizational theory: institutionalized
trust. And it has a delivery mechanism: communication, first spoken, then
written, then encoded into the procedures, contracts, and bureaucracies that
let a stranger in Rotterdam trust a signature from a stranger in Chennai.
Bureaucracy earns a bad reputation as red tape, but at its root it is a trust
technology, a set of standardized procedures that lets unrelated people
transact without needing to personally verify each other’s character.
Double-entry bookkeeping, the joint-stock company, the postal system, the
passport, these are all inventions in the same category as language itself:
tools that let cooperation scale past the roughly 150 people (Robin Dunbar’s
famous number) that our brains can track through personal relationship alone.
Physical technology has always ridden alongside this social technology,
amplifying it. The wheel didn’t just move goods faster, combined with
standardized axle widths and road networks, it made regional trust networks
viable at distances no courier on foot could sustain. The printing press didn’t
just reproduce text, by making the
Bible, and later pamphlets, newspapers, and scientific journals, available to
anyone literate, it broke the monopoly that scribal elites held over what
counted as agreed-upon truth, and in doing so it re-founded whole religious and
political orders. The telegraph, the telephone, the internet: each one is,
underneath the marketing, a new trust-and-coordination layer stacked on the
ones before it.
Now comes artificial intelligence, and it is not just another rung on
that ladder. It is different in kind, because for the first time the tool
doesn’t merely transmit human-generated trust signals faster , it can generate judgment, synthesis, and
decisions on its own. That changes the question. It’s no longer just “how fast
can strangers coordinate” but “who gets to do the coordinating, and on whose
behalf.”
There are two directions this can go, and they lead to very different
civilizations.
The first direction: the
Borg model
A small number of frontier labs - right now, realistically, a handful of
companies in two countries - train models on a scale of data and compute that
nobody else can replicate: the entire searchable internet, increasingly private
data through partnerships and acquisitions, and enough proprietary usage logs
from hundreds of millions of daily conversations to know how people think,
argue, and decide better than the people know themselves. Everyone else becomes
a client, querying a central intelligence that has never revealed what it
learned about them to anyone but itself. It is not an accident that the Star
Trek Borg is the right metaphor: a distributed set of drones, individually
unremarkable, whose intelligence is aggregated upward into a Queen who alone
sees the whole picture and alone decides. Assimilation doesn’t require malice.
It only requires that everyone’s data flows one way upwards , while judgment flows back down as a service.
The second direction:
diffusion
Instead of one model trained on everyone’s data and queried by everyone,
imagine a nested architecture of intelligence that mirrors the way trust itself
has always scaled in human societies - from individual to family to community
to nation, each layer adding coordination without fully surrendering what came
before it. A personal model that learns primarily from an individual’s own
history and stays substantially theirs. A household-level router that
reconciles the family’s shared needs, finances, schedules, health, without exporting
the raw data to any central party. Community and institutional layers that pool
just enough signal to coordinate like a hospital network sharing anonymized
treatment outcomes, a farming cooperative sharing yield data and so on without
surrendering the underlying record. National or civilizational layers that
federate further still, for the genuinely public-goods problems: pandemic
response, climate modeling, financial stability. Intelligence increases with
altitude, but so does the friction required to extract raw data upward. This is
closer to how evolution actually organizes complexity - through modular, semi-autonomous units that
coordinate without fully centralizing control - than the single-brain model the
Borg represents.
The diffusion model is the one worth betting on, but it is worth being
honest about why it isn’t automatic. Model weights being “open” or a chatbot
running locally on a phone does not, by itself, redistribute power. Underneath
even the most local-feeling AI product today sits a stack that is still
extremely concentrated: pretraining compute that only a few labs can afford,
chip design and fabrication controlled by a handful of firms in a handful of
countries, and energy infrastructure that is itself a scarce, geopolitically
contested resource. An open-weight model trained on a closed, centrally-scraped
corpus is diffusion in name and centralization in substance , a longer, more
comfortable road to the same Queen. If this century’s version of the printing
press turns out to require a printing press factory that only three governments
can build, the diffusion story collapses into the Borg story with better
marketing.
There is, however, an answer to the training problem, and it comes from
the closest analogy available: how humans themselves acquire capability. Every
human being is “pretrained” on a broad common corpus - language, schooling, the
accumulated knowledge of a culture - before developing anything distinctive.
Universal education does not centralize human intelligence; it equips each mind
to then learn recursively from its own experience, in directions no curriculum
planned. The base model can play the same role: a common endowment, trained
once on broad public data, the way a public education system is funded once for
everyone. Diffusion becomes real at the point past that endowment , when each
node in the hierarchy, whether an individual, a household, a firm, or a
community, holds not merely a copy of the model but the capacity to keep
learning from what it alone can see, and when the owner of that node decides
what portion of the learning is exposed upward. This mirrors how capability has
always worked in human society. A doctor shares her diagnosis, not the decade
of pattern recognition behind it; a firm sells its product, not its process
knowledge; a family teaches its children things it would never publish. Skill
and disclosure have always been separable, and that separability is precisely
what a query-everything-through-the-center architecture destroys - the center
learns from every interaction, while the individual accumulates nothing that is
durably theirs.
Honesty requires admitting that this recursive-learning-at-the-edge
capability does not fully exist yet. Most of what is marketed today as
personalization is retrieval: the model consults an individual’s documents and
history at query time without changing itself, which means the accumulated
learning still lives wherever the model lives. Genuine local learning, models that update themselves from experience,
cheaply, on modest hardware, without catastrophically forgetting what they
already knew remains a hard, open engineering problem. The technical
trajectory, to be fair, is bending in the right direction: models keep getting
smaller for a given level of capability, techniques for cheap adaptation keep
improving, and consumer chips now ship with dedicated neural hardware as a
matter of course. What is not bending is the commercial trajectory. The
economics of every frontier lab reward keeping the learning loop at the center,
because centrally accumulated learning is the moat, the more the central model
learns from everyone, the harder it becomes for anyone to leave. So the two
trajectories diverge: feasibility is diffusing outward while deployment keeps
concentrating inward, and it is precisely in that gap that policy has work to
do. Recursive learning at the edge will not be handed down by incumbents whose
business model it undermines; waiting for the market to deliver it is like
waiting for the scribes to distribute the printing press. It has to be pulled
forward deliberately by public research
funding, by procurement rules that require publicly purchased AI systems to
support local learning and owner-controlled disclosure, and by writing the
principle that the learning stays with the learner into data protection law,
the way purpose limitation was written in a generation ago.
So the honest version of the bet is not “small models will save us.” It
is that diffusion has to be built deliberately, at every layer of the stack,
the way earlier trust infrastructure was built deliberately, through standards,
law, and public investment, not left to emerge on its own from a market that
has every incentive to concentrate.
There is precedent for exactly this kind of deliberate construction, and
it is worth pointing to because it already exists rather than remaining
hypothetical. India’s approach to digital public infrastructure - a unified
payments protocol that any bank or fintech can plug into rather than routing
transactions through a single dominant platform, a verifiable-credentials
system that lets individuals hold and share their own documents rather than
surrendering them to a central database, and an open commerce network that lets
buyers and sellers transact across competing apps rather than being locked into
whichever platform got there first, is
essentially an attempt to build trust infrastructure as a shared, low-lock-in
utility rather than as proprietary rails owned by one company. It is not a
perfect model and it has real gaps, but it demonstrates something important:
that population-scale coordination doesn’t require a single controlling entity
if the protocol layer is deliberately kept open and interoperable. The lesson
for AI is not “copy this system” so much as “copy the design principle”, build
the equivalent of open rails for identity, data portability, and model access,
so that intelligence can be composed from below rather than only distributed
from above.
What would that take in practice? A few concrete interventions, none of
them exotic:
•
Data portability as an enforceable right, not a feature,
so an individual’s interaction history
can move with them between AI providers the way a phone number now moves
between carriers, preventing lock-in from doing quietly what outright control
could not do openly.
•
Public or multilateral investment in compute and energy
capacity outside the two or three countries that currently dominate it, on the
model of how public investment built highways and rural electrification rather
than waiting for private markets to reach unprofitable places on their own
schedule.
•
Interoperability standards for model-to-model and
agent-to-agent communication, so a household-level or community-level system
can coordinate with a national one without needing to be owned by the same
company that owns the national one, the
AI equivalent of the postal system agreeing on envelope sizes.
•
Regulatory pressure specifically aimed at the
infrastructure layer, chips, cloud
capacity, energy contracts , rather than only at the visible chatbot layer,
since that is where real concentration risk is currently accumulating fastest
and most invisibly.
•
A cultural shift among the capable middle tier of
nations, those with talent, institutions, and ambition but not
frontier-lab-scale capital, toward building shared, federated capability with
each other rather than each negotiating bilaterally and separately with the
handful of dominant labs, which only reproduces a hub-and-spoke Borg structure
one client relationship at a time.
Sceptics of the diffusion path will point out, correctly, that some
problems genuinely need a Queen, or at least a very large brain. Pandemic modelling,
climate prediction, and fundamental scientific discovery benefit from the kind
of massive, centralized compute that only a handful of institutions can field,
and no household-level router is going to fold a protein or model a hurricane.
The diffusion argument is not that centralized capability should not exist. It
is that centralized capability should be treated the way we treat other
infrastructure with natural concentration risk, nuclear power, undersea cables,
the electrical grid, as a public utility
subject to oversight, access rules, and accountability, rather than as the
private property of whichever company got there first. The European Union’s AI
Act, whatever its flaws in execution, is at least an attempt to draw that line:
to say that as models approach systemic scale, the obligations on their
operators should scale with them. Export controls on advanced chips are a
cruder version of the same instinct, aimed at slowing the concentration of the
compute layer rather than the software layer, even if their current form is
more about geopolitical rivalry than about distributing power more broadly.
And the early scaffolding for genuine diffusion is already visible.
Federated learning in healthcare, hospitals training shared diagnostic models
by exchanging model updates rather than patient records, demonstrates that collective intelligence does
not require pooling raw data in one place. On-device inference, now standard on
flagship phones, means a growing share of everyday AI use never has to leave
the device at all. Neither fully solves the concentration problems described
above, but both show that the direction is technically viable; what is missing
is the institutional will to deploy such architectures at population scale
rather than leaving them as premium features for those who can already afford
to ask.
Humanity’s edge was never raw intelligence. It was the invention of trust
technologies that let intelligence combine across strangers without requiring a
single mind to hold it all. Writing, law, currency, and bureaucracy did this by
distributing judgment while standardizing the interface between people. Whether
AI becomes a fifth trust technology in that lineage, or the tool that finally
lets a single mind hold it all, is not a question that resolves itself as
models get better. It resolves according to who builds the rails underneath
them, and how deliberately the rest of us insist that those rails stay open.
That is a choice still being made, right now, mostly in rooms far from public
view — which is exactly why it needs to be argued for in public.
"We taught the whole species to read.
We did not hand every book to one reader."
Saturday, July 11, 2026
India’s Innovation Strategy and the China Misread
India is assembling an industrial-policy toolkit that includes
production-linked incentives, the IndiaAI Mission, semiconductor subsidies, and
lessons drawn from global innovation systems. The instinct is understandable.
Governments want to compress technological catch-up through coordination and
capital. Yet the lesson India appears to be drawing is more complicated than
either its admirers or critics suggest.
The most consequential Chinese technology outcome of this decade was not
produced by the Chinese state in the way it is often assumed. DeepSeek, the AI
firm whose low-cost frontier models unsettled Silicon Valley and reshaped
assumptions about the price of intelligence, did not emerge from a national
champion program. It was not a state-picked winner under a five-year plan. It
did not originate inside China’s formal industrial policy machinery.
India’s Innovation Strategy and the China Misread
Click to read on the full article published in Transcontinental Times
Sunday, June 7, 2026
AI Governance and Future of Work
My Speech
at AI-DPI – 26 Conference organised by NCEAR
Let me begin with a simple observation that I think frames everything we're about to discuss.
In the last
few centuries, we witnessed multiple technological disruptions ranging from
printing press to industrial revolution to computers to internet. It restructured
society what work meant, where people lived, what skills had value, what
governments needed to do. The economic and social ripple effects played out
over decades.
Today, we
are living through a transformation that is much more profound but the ripples
are moving in months, not decades. And that gap between the speed of
technological change and the capacity of our institutions to respond is
precisely why conversations like this one matter.
Welcome to
what I hope will be a frank, insightful, and perhaps uncomfortable conversation
about AI, governance, and the future of work.
In the last
three years, artificial intelligence has crossed a threshold that surprised
even its creators. Large language models can now draft contracts, write code,
analyse medical scans, counsel customers, generate creative content, and
conduct research, tasks that, until recently, defined the upper tier of
knowledge work.
We are no
longer talking about AI that automates the routine. We are talking about
AI that can perform the cognitive. That is a qualitatively different
kind of disruption, and it demands a qualitatively different kind of response.
Three
tensions sit at the heart of today's discussion.
First tension
is on Governance
When it
comes to governance of AI key questions that arise are
Who governs
AI? Who benefits? Who bears the cost of disruption? These are political and
moral questions, not just technical ones.
There lies the
tension between speed and safety. AI development is moving at a pace
that regulatory frameworks were simply not designed to match. The EU AI Act
took years to negotiate and is already facing questions about whether its risk
categories reflect the technology as it exists today, let alone as it
will exist in three years. India is developing its own digital governance
frameworks, and the choices made here, given the scale of this country's
workforce and its digital ambitions, will matter not just domestically but
globally.
The core
challenge for governance is this: if you regulate too slowly, you cede the
field to actors, corporate or national, who face no constraints. If you
regulate too quickly, you risk encoding today's assumptions into law and
stifling the innovation that could actually solve problems. There is no
comfortable middle ground. There is only the hard work of trying to get it roughly
right, fast enough to matter.
Here the tension
is also between innovation and accountability. The companies building
the most powerful AI systems are, understandably, advocates for an environment
that allows rapid development. Many also, to their credit, genuinely grapple
with questions of safety and responsibility. But the incentive structures of
competitive markets are not naturally aligned with the kind of careful,
transparent, accountable development that the stakes of this technology
require.
Governance,
at its best, creates the conditions under which accountability becomes not a
constraint on innovation but a foundation for the trust that allows
innovation to scale. We do not have that governance architecture yet. Building
it nationally and internationally is one of the defining challenges of this
decade.
Next tension
in in he "Future of Work" that is Already Here
There are three
competing narratives on this paradigm shift
- Displacement: AI replaces human jobs at
scale
- Augmentation: AI makes workers more
productive and valuable
- Transformation: New categories of work emerge
that we can't yet name
Here lies
the tension between productivity and dignity. Every
study that examines AI's impact on knowledge work shows significant
productivity gains. Legal researchers, coders, financial analysts, writers when
well-supported by AI tools, they produce more, faster, and often at higher
quality. This is genuinely good news.
But
productivity gains do not automatically translate into widely shared
prosperity. The history of technological disruption is also a history of
transition costs borne
disproportionately by workers who lack the resources, the retraining
opportunities, or the institutional support to adapt. The question is not
whether AI will transform work. It will. The question is whether that
transformation will be something we navigate together or something that happens
to millions of people who had no voice in shaping it.
India is not
a passive observer in this story. It is one of the central actors.
This country
has one of the world's youngest and most rapidly digitising workforces. It has
a technology sector that has spent decades building the global knowledge
economy's operational backbone. It has a government that has shown real
ambition in digital public infrastructure, from UPI to Aadhaar to the Open
Network for Digital Commerce.
And it faces
a specific, urgent challenge. A significant proportion of India's IT and BPO
workforce, millions of skilled, educated, middle-class workers are employed in
precisely the categories of knowledge work that generative AI most directly
disrupts. Customer support, document processing, software testing, data
annotation, back-office operations. These jobs are not going away tomorrow. But
the trajectory is clear, and the window for preparation is not infinite.
At the same
time, India has something that not every country has in this moment: scale
as an asset. The diversity and volume of India's linguistic, cultural, and
domain-specific data; the depth of its technical talent; its position as a
potential standard-setter for the Global South in AI governance, these are
genuine opportunities, if they are seized with intention.
I'll offer
three propositions to anchor our panel discussion.
First: governance
must be adaptive, not just reactive. We need regulatory frameworks that are
designed to evolve, that build in review cycles, that involve multistakeholder
input, that distinguish between the risks of different applications rather than
treating AI as a monolithic category. A diagnostic AI in a hospital has
different risk parameters than a recommendation algorithm on a social platform.
Governance that treats them identically will either over-regulate the
beneficial or under-regulate the harmful.
Second: the
future of work requires active investment, not just passive optimism. It is
not enough to say that new technologies create new jobs, historically, they
often do. What matters is the transition: whether workers have access to
retraining, whether institutions like schools and universities adapt their
curricula in time, whether social safety nets are designed for an economy where
the nature of employment is changing. This is a policy challenge, not just a
market one.
Third: the
voices in the room must expand. The conversations that shape AI governance
tend to happen in a relatively small number of rooms, boardrooms, regulatory
agencies, international standards bodies, academic conferences. The people
whose working lives will be most directly transformed are rarely in those
rooms. That needs to change , not as a matter of procedural fairness, but
because the decisions will be better if the inputs are broader.
Let me
close with this.
I am neither
a pessimist nor an optimist about AI. I am a realist who believes that the
outcomes of this transformation are genuinely open that they will be determined
not by the technology alone, but by the choices we make about how to develop
it, deploy it, govern it, and distribute its benefits.
The future
of work is not written. It is being written right now, in the decisions being
made in companies, in legislatures, in classrooms, and in conversations like
this one.
My hope for
today is that we leave this room with sharper questions, not just comfortable answers
and perhaps with a clearer sense of where action, not just analysis, is
required.
Thursday, May 28, 2026
AI, Costs, and the Myth of Inevitable Human Obsolescence
Why
the disruption narrative is more complicated, and more hopeful, than it appears
For years, the AI narrative has been relentlessly linear and tinged with
apocalypse: models get smarter, cheaper, and more capable, and humans get edged
aside, role by role, sector by sector. Then came a headline that disrupted the
script.
“Microsoft is limiting internal use of expensive AI coding tools as
enterprise AI costs surge.”
It sounded like a contradiction from the company that bet its future on
AI, poured $80 billion into data centres, and plastered Copilot across every
product it makes. It was not a contradiction. It was a revelation, though, as
we shall see, a more layered one than the headline suggests.
The Economics Behind the Curtain
Inside Microsoft’s engineering divisions, Claude Code, the AI coding
assistant from Anthropic, was not cancelled because it failed. It was cancelled
because it succeeded too well. Rolled out to roughly 5,000 engineers in the
division behind Windows, Microsoft 365, Outlook, Teams and Surface, it reached
usage rates of 84 to 95 percent within months. Engineers used it relentlessly,
and token-based billing, where every prompt, every agentic loop, every
code-generation cycle costs real money, ran to an estimated $500 to $2,000 per
engineer per month. The internal memo set a cancellation deadline of June 30,
2026.
The pattern is wider than Microsoft. Uber’s CTO has confirmed the company
burned through its entire planned 2026 AI coding budget in four months, after
actively incentivising engineers to maximise usage. Meta built an internal
leaderboard called “Claudeonomics” to track which employees were consuming the
most AI tokens. Amazon encouraged “tokenmaxxing”, gamifying maximum AI
consumption as a proxy for productivity.
Two honest caveats belong in this story, and most commentary has skipped
both.
First, Microsoft’s decision was not purely about cost. Engineers
reportedly preferred Anthropic’s tool to Microsoft’s own Copilot CLI, and the
cancellation conveniently redirects them into Microsoft’s own stack. Cost was
real; so was competitive strategy. Second, the escape route is no escape:
GitHub Copilot itself is moving to usage-based billing from June 2026. The
token meter is not a Claude problem. It is the emerging price structure of
frontier AI itself.
The collective result stands nonetheless: for many enterprise tasks
today, undisciplined AI usage is more expensive than the humans it was meant to
augment. The promise was frictionless efficiency. The reality is that when
thousands of employees use frontier models without governance, the economics
invert.
The Objection This Argument Must Survive
Before drawing conclusions, the cost story has to face its strongest
counter-argument: the price of intelligence is falling, fast. The cost of
frontier-quality inference has been dropping several-fold every year. What
looks like a “cost ceiling” in 2026 could look like a rounding error by 2028.
So is the Microsoft episode just a temporary blip?
Not quite, and the reason is an old one. Economists call it the Jevons
effect: when something useful gets cheaper, we do not spend less on it; we use
vastly more of it. Microsoft’s engineers did not hit a budget wall because
tokens are expensive. They hit it because usage exploded faster than prices
fell, agentic loops, always-on assistants, code generated and regenerated at
industrial scale. Every cost decline to date has been swallowed by appetite.
The lesson is not that AI is permanently expensive. The lesson is that AI
consumption, like cloud computing before it, will always expand to consume the
budget available, and therefore governance of usage, not the price of tokens,
is the durable management problem.
Does This Slow the Replacement of Humans?
Only partially, and only temporarily. The cost ceiling buys time. It does
not change direction.
The “AI will replace humans” narrative was always too blunt. AI is not a
flat substitute for human labour. It has a cost curve. At low usage it is
remarkable. At scale, without discipline, it is ruinous. Companies will not
replace human beings wholesale. They will replace them selectively, deploying
AI where ROI is unambiguous, retaining humans where judgment, accountability,
ambiguity, or trust cannot be priced away.
The most exposed roles are not at the bottom of the skills ladder or the
top. They are in the middle: paralegals, junior coders, financial analysts,
content writers, customer service agents, roles that are routine,
pattern-based, and high-volume. At Uber, around 70 percent of committed code
now originates with AI. That number should concentrate the mind of every
mid-career professional whose work is pattern recognition at volume.
The Radiology Test: Why “Exposed” Is the Wrong Word
But “exposed” is a one-dimensional lens, and one profession shows why.
Consider radiology, the example most often cited, for a decade now, as the
first white-collar casualty of AI.
On capability, the pessimists are right: AI can read many scans as well
as or better than humans, and the per-unit economics are unanswerable. On
accountability, the pessimists are early: in most jurisdictions, a diagnosis
requires a licensed human signature, and regulators, courts and insurers will
keep it that way for some time, not because the human is always more accurate,
but because someone must be answerable when the machine is wrong.
And on access, the pessimists have the sign of the effect backwards, at
least in a country like India. The binding constraint on radiology in
small-town and rural India has never been an oversupply of radiologists. It is
that there are almost none. AI-assisted reading changes that arithmetic. A
local physician in a taluk hospital, supported by AI triage and a remote human
radiologist for sign-off, can now order and act on imaging that was previously
out of reach. The realistic effect in such markets is not fewer radiology jobs
but more radiology, more referrals, more scans, more diagnostic activity, and
new paramedical and technician roles around it, in places where the alternative
was not a human radiologist but nothing at all.
The same technology, in the same year, can displace work in saturated
markets and create it in underserved ones. Radiology is not an exception; it is
the template. Apply the same three questions, can AI do it, who must answer for
it, and where was the service never available at all, to law, to accounting, to
software, to education, and the picture that emerges is not a single wave of
obsolescence but a redistribution: of tasks within professions, and of services
across geographies.
The Real Disruptor: Robotics, Not Software
While enterprises wrestle with token bills, robotics companies are
solving a different equation. A humanoid robot is largely a capital
expenditure: its “salary” is electricity, maintenance and software. Tesla
Optimus, Figure and Boston Dynamics are targeting price points intended to
undercut the minimum wage in developed economies within this decade, beginning
with exactly the jobs that employ hundreds of millions globally, fast food,
warehouse picking, hotel housekeeping, retail stocking.
Two qualifications keep this honest. First, the cost structure is
different from software AI, not free of recurring costs: many robotics firms
are pricing robots-as-a-service, with subscriptions and teleoperation support,
a cousin of the token meter, not its opposite. Second, the crossover point
depends on the wage it must undercut. A robot that beats a $15-an-hour wage in
Ohio is nowhere near beating a ₹15,000-a-month wage in Kanpur. Which leads to a
striking inversion: the Global South will likely adopt software AI fastest,
because models are cheap and skilled labour is scarce, and adopt robotics
slowest, because physical labour is abundant and cheap. The displacement map of
the next decade will not be uniform. It will be a patchwork drawn by local
wages.
The China Factor: The Cost Floor May Collapse
DeepSeek’s breakthrough in early 2025, and the rapid succession of
low-cost Chinese models since, changed the global cost equation in ways that
have not been fully absorbed. Frontier-level reasoning delivered at a fraction
of Western pricing, backed by structural advantages: subsidised compute and
energy, lower infrastructure costs, and less shareholder pressure to monetise
quickly.
If such models gain wide adoption across India, Southeast Asia, Latin
America and Africa, markets where price matters more than geopolitics, the cost
barrier falls years ahead of current projections. And note that this cuts both
ways: cheap models accelerate displacement of routine work, but they equally
accelerate the access story told above. The same collapsing cost floor that
threatens the call-centre agent makes the AI-assisted rural clinic viable. The
West will move more cautiously, constrained by data sovereignty and regulatory
anxiety. The Global South may move faster, precisely because it cannot afford
to be slow.
This creates a two-speed world of AI adoption, and, by extension, a
two-speed world of both displacement and inclusion. That asymmetry deserves far
more attention than it receives in the global policy conversation, which
remains written almost entirely from the vantage point of high-wage economies.
The Jobs That Don’t Exist Yet
Every major technological shift destroys familiar work and creates
categories of work invisible until they become indispensable. The steam engine
did not just displace handloom weavers; it created railway engineers and
factory inspectors. The internet did not merely kill travel agencies; it
created cloud architects and UX designers. The honest difficulty is that new
jobs are not legible until they exist.
Still, the outlines are forming at the margins: people who design how
humans and AI divide work and supervise each other; specialists who generate
and govern synthetic data as real data becomes regulated and scarce;
professionals who arbitrate between AI outputs and human decisions in finance,
healthcare and law; a new blue-collar profession maintaining and supervising
robot fleets; AI safety and governance analysts, a field that will grow to the
scale of cybersecurity; and millions of AI-enabled one-person enterprises that
would have been operationally impossible a decade ago. Above all, as AI absorbs
the burden of logic and pattern, the premium on trust, empathy, cultural
context and meaning rises, precisely because AI cannot credibly supply them.
AI as an Equaliser: The Welfare Dimension
If AI dramatically reduces the cost of delivering essential services, the
welfare gains could, under the right policy conditions, outweigh the disruption
to employment. AI-optimised irrigation and supply-chain prediction could raise
yields and cut food waste across the Global South. AI triage, diagnostic
support and remote monitoring could bring quality care to populations who currently
have none, the radiology story above, repeated across a dozen specialities.
Adaptive AI tutors could personalise learning for hundreds of millions of
children for whom quality schooling remains a geographic accident of birth.
If AI reduces the cost of food, health and education by half or more, the
question is no longer simply “who loses their job?” It becomes “what kind of
society do we build with the surplus?” That is a question of political will,
not technology.
Where Humans Still Win on ROI
Despite the compression of timelines, some domains will remain higher-ROI
for humans through this decade: trust-based, relationship-driven work, where
clients pay a premium for human accountability; novel problem-solving in
genuinely ambiguous environments, where no training data exists for the
situation at hand; skilled trades in unstructured physical settings, the
plumber navigating an unfamiliar home, the electrician improvising under
deadline; and regulated roles where law or professional standards mandate human
sign-off regardless of AI capability.
The pattern across these safe harbours is consistent. What makes humans
irreplaceable is not intelligence alone. It is accountability, physical
adaptability, and the fact that in some relationships the human presence is
the product, not merely the mechanism of its delivery.
The Question That Actually Matters
The debate has moved on from whether AI will displace human workers. The
debate now is: which humans, doing which tasks, in which geographies, on what
timeline, under which cost structures, and, critically, will the new categories
of work emerge quickly enough, and be accessible enough, to absorb those
displaced?
The cost ceiling revealed by Microsoft, Uber and others is real. It slows
the slope, forces selectivity, and creates space for societies to adapt rather
than absorb a vertical shock. But it does not alter the destination. The
direction of travel has not changed, only the gradient of the curve. And the
gradient, as the radiology test shows, points in different directions in
different places: downward for routine work in saturated markets, upward for
services in markets that never had them.
The most important investment any individual, institution, or government
can make right now is not in AI itself. It is in the human capacity to navigate
the transition: to identify which skills will compound in value, which roles
are building toward the new categories, and which paths are quietly narrowing.
“AI will not erase human value. It will redraw the map of where that value lives, and whether we prepare to inhabit that new terrain is the defining challenge of this decade.”
Sources:
The Verge (Tom Warren, Notepad, May 14, 2026) on Microsoft’s internal Claude
Code cancellation; The Information (April 2026) on Uber’s AI coding budget;
Fortune (May 22, 2026) on Meta’s “Claudeonomics” and Amazon’s token incentives.
Friday, May 8, 2026
E = MC² : The Equation That Never Gets Old
On Measurement, Continuous Improvement, and Customer Focus — Then and Now
There is a particular kind of excitement that technology
companies are exceptionally good at, and a particular kind of discipline they
are chronically bad at. The excitement is building. The discipline is running.
Every new feature, every new product, every new platform gets showered with
energy, talent, and attention. The unglamorous work of making sure it all
actually works, consistently, reliably, at scale, day after day, gets left to
whoever is available, measured by whatever is easy to measure, and improved
only when something breaks badly enough to be embarrassing.
This is not a new observation. But it has become a vastly
more consequential one. Because we are now deploying AI systems and autonomous
agents into operational environments at a pace that far outstrips our
willingness, or our ability, to govern them. And the cost of that gap is no
longer measured in minor inefficiencies. It is measured in compounding,
invisible failures, in decisions that
are wrong by design, in resources consumed by systems nobody is watching, and
in customers quietly harmed by processes nobody is truly accountable for.
The answer to this is not more technology. It is better
operational discipline. And the framework for that discipline is simpler than
most people think.
We call it E = MC²: Excellence, derived from a
culture that Measures relentlessly, pursues Continuous
improvement, and never loses sight of Customer focus. These three
elements are not independent. They are a virtuous cycle, each one feeding the
others, each one incomplete without the others. Understanding how they connect,
and how to make them real, is the central challenge of operational management
in any era. Including this one.
Why Measurement is Hard, Even for People Who Handle Data
for a Living
There is a paradox at the heart of the IT and services
industries. These are sectors whose entire value proposition rests on data, on
capturing it, organising it, analysing it, and making it useful. And yet, in
practice, their internal operational measurement discipline is often
surprisingly immature. The processes that organisations build for their
customers are rarely applied with equal rigour to their own operations.
The reasons are not mysterious. The glamour in these
industries flows toward novelty, toward "cool functions,"
"exciting features," and "latest gadgets." Boring pursuit
of efficiency gains simply does not compete for talent or attention. When a
senior engineer has a choice between building something new and spending six
weeks instrumenting something old to understand why it sometimes fails, the
outcome is predictable. And so operational measurement tends to happen
reactively, in response to a crisis, a
customer complaint, or a regulator's inquiry, rather than as a continuous, proactive
discipline.
To learn how to do this differently, it helps to look at
industries that never had the luxury of treating operations as an afterthought.
The hazardous chemical process industry is an instructive
model, and not an intuitive one. It has been around for centuries, long enough
to have matured its operational practices through hard experience. Its product
lines are largely commoditised, which means margins are thin and efficiency is
not optional, It is existential. The consequences of process failures are
sometimes fatal, which means the scrutiny. public, regulatory, and internal, is
unrelenting. And its processes are integrated end-to-end, with limited
visibility into what is actually happening inside the pipes at any given
moment, which forces a culture of strong monitoring and control.
These are, in fact, exactly the conditions that characterise
complex digital operations today. Thin margins. High stakes. Limited internal
visibility. Regulatory scrutiny. The main difference is that the chemical
industry has spent decades building the measurement culture to match these
conditions, while the technology and services industries are still, in many
cases, at the beginning of that journey.
From that more mature tradition, three elements of
measurement discipline emerge as foundational.
The Three Pillars of Measurement
Flow Management: Count Every Transaction
The first pillar is what might be called micromanagement of
the operation, not in the pejorative sense of hovering over people, but in the
precise sense of tracking each input through each sub-process it was meant to
traverse, confirming it arrived correctly and without error.
This sounds obvious. In practice, it is done poorly, or not
at all, especially for processes that are still evolving. When a new system or
workflow is still being refined, exceptions proliferate. And exceptions, in
young computerised systems, have a dangerous tendency to become invisible, swallowed by automated retry mechanisms,
silently skipped, or classified as edge cases that never quite make it onto
anyone's priority list.
The consequences of poor flow management are almost always
financial and reputational, and they tend to be discovered embarrassingly late.
A large bank once sent letters to its credit card customers admitting that it
had not been tracking transactions correctly, and asking recipients to settle
on the basis of their own personal records. The transactions were not hidden.
They were not stolen. They had simply not been tracked. The systems were
running; the accounting was not. When providers of transaction billing
solutions are brought into organisations for the first time, the revenue
leakage they surface, from transactions that fell through the cracks of
inadequately monitored processes, is routinely staggering.
These are not exotic failures. They are the entirely
predictable consequence of building systems without building the measurement
infrastructure to watch over them.
Capacity Management: Know Where the Bottlenecks Are
Before They Happen
The second pillar is the macro view, tracking the capacity
of processes, people, service providers, and machines in order to anticipate
bottlenecks before they become crises. This requires establishing trend
measures for each element and monitoring them continuously, not just
periodically.
Capacity management is especially treacherous in
computerised environments for a structural reason: shared resources. Network
infrastructure, compute capacity, database connections, these are all consumed by multiple processes
simultaneously, and the utilisation curve for each process grows differently. A
system that appears to have adequate capacity for today's workload may have
none for tomorrow's if the growth curves are not being watched and modelled.
Two particular categories of hidden capacity consumers
deserve special attention, because they are pervasive and almost universally
underestimated.
The first is queries. Every business generates a need for
data extracts — for management reporting, regulatory compliance, customer
service lookups, and ad hoc analysis. These queries consume the same production
capacity as the operational processes. And they are disproportionately likely
to be written inefficiently, because they are typically assigned to junior
resources or business analysts who lack the training to optimise them, and
because there is very little accountability for query performance until something
breaks. A query that was meant to run once becomes a standard report. A
standard report that runs nightly becomes a standard report that runs hourly.
The cumulative resource consumption creeps upward invisibly until, one day, the
system slows to a crawl during peak operational hours, and nobody can
immediately explain why.
The second is design debt. For most software developers, the
genuine satisfaction is in building features. Once a feature is live and
functioning, interest moves on. The pressure to optimise, to refactor, to
improve efficiency, runs directly against the incentive to ship the next thing.
The result is that bespoke systems accumulate performance inefficiencies that
are never addressed , not because fixing them is technically difficult, but
because nobody is measuring the cost of leaving them in place, and nobody is
accountable for the cumulative drag. In most organisations, there is scope for
at least a hundred percent improvement in process efficiency simply by
addressing the worst of these design inefficiencies, but only if someone is measuring for them.
Service Levels: Commit to the Customer, Then Track the
Commitment
The third pillar is where measurement connects most directly
to purpose. The most powerful mechanism for ensuring that measurement and
improvement activity stays focused and meaningful is to define, publicly and
clearly, what the organisation is actually committing to deliver to its
customers.
There is an important distinction to draw here between a
Service Level Agreement and what might be called a Customer Service Commitment.
An SLA is a floor — a formal definition of the minimum below which the
organisation will try not to fall. It is a legal and contractual instrument,
and it tends to create a culture of adequacy: as long as we are above the
floor, we are fine. A Customer Service Commitment is something different. It is
a genuine aspiration — a statement of what the organisation sincerely believes
it can and should deliver, at a level meaningfully above the minimum.
This distinction matters because people and systems tend to
optimise for what they are measured against. An organisation that measures
against its SLAs will manage its operations to the SLA threshold. An
organisation that measures against its Customer Service Commitments will manage
its operations to the standard it actually believes in.
The mechanics of tracking these commitments deserve specific
attention. Time-series data, tracking key performance parameters not just at a
point in time, but continuously over time. is essential for detecting trends
before they become crises. A single data point t ells you where you are today.
A trend tells you where you are going. And it is the trend that matters
operationally, because by the time a single bad reading turns into an obvious
crisis, the window for preventive action has usually closed.
It is also worth having the team that tracks customer
commitments sit separately from the team responsible for operations. This is
not about distrust. It is about the structural reality that an operations team
under pressure will, understandably, interpret ambiguous data in the most
favourable light available. A separate tracking function provides the
independent visibility that makes measurement honest.
Continuous Improvement: From Counting to Acting
All of this measurement serves one purpose: enabling the
organisation to improve, continuously, before it is forced to by failure.
This is more difficult than it sounds, because the culture
required to use data for continuous improvement is fundamentally different from
the culture most organisations actually have. In most places, data tracking
reports are either compliance artifacts — produced to satisfy an audit or a
boss, or post-mortem instruments, pulled out after something has gone wrong to
explain what happened. Neither of these uses generates improvement. They
generate paper trails.
The culture of continuous improvement requires something
harder: the regular, disciplined use of data to find problems that have not yet
caused visible failures. This means looking at trend shifts before they become
obvious. It means investigating unusual volatility in metrics that are still
technically within acceptable bounds. It means preferring prevention over
heroism — which runs directly against the organisational instinct that rewards
the person who fixed the crisis rather than the person who avoided it.
To make this a habit rather than an occasional initiative,
it has to become a ritual. The cadence of reviewing operational data,
identifying trends, assigning root cause investigations, and tracking
improvement actions has to be embedded into the organisation's regular rhythm,
not treated as an additional burden on top of "real work." When it is
done well, it does not feel like overhead. It feels like the organisation
learning from itself in real time.
The AI Era Changes the Stakes, Not the Principles
Everything described above was relevant in 2009. It is more
relevant now by an order of magnitude.
The introduction of AI systems and autonomous agents into
operational environments does not render these principles obsolete. It makes
them urgent. Because AI introduces a new category of operational actor, one
that is more capable, more opaque, and more consequential than anything that
preceded it, into environments that, in
many cases, barely had adequate measurement cultures to begin with.
The most important thing to understand about AI in
operations is that it fails in ways that are qualitatively different from how
conventional software fails. Traditional software fails visibly. A system
crashes. A transaction errors out. A service goes down. These failures are, in
their own way, manageable, because they announce themselves. AI fails silently.
A model that has drifted from its training data continues to generate outputs
that look confident and coherent, while producing decisions that are subtly,
systematically wrong. A recommendation engine with a bias baked into its
training data does not flag an anomaly; it just consistently disadvantages
certain customers. A document processing agent that hallucinates does not throw
an exception; it produces a confident, plausible, and incorrect result.
This is the flow management problem, rewritten for the age
of AI. Every AI-powered process needs a systematic accounting not just of what
it produces, but of the quality, reliability, and drift of those outputs over
time. The input went in; the output came out, but was the agent's reasoning
within acceptable bounds? Was its confidence calibrated? Were there exceptions
that the system silently swallowed rather than escalating to a human? The
revenue leakage and customer harm that flow from unmonitored AI processes make
the untracked credit card transactions of an earlier era look quaint.
The capacity management problem is also fundamentally
transformed. AI models are the most resource-intensive entities ever introduced
into enterprise operations. A single large model inference can consume more
compute than an entire legacy application stack, and when multiple agents run
concurrently, as they increasingly do, in agentic architectures where AI
systems orchestrate other AI systems, the shared infrastructure constraints
become genuinely complex to manage. The hidden capacity consumers have
multiplied: poorly designed prompts that generate verbose, expensive outputs;
inefficient agent chains that make redundant calls; one-time AI automations
that quietly become permanent fixtures eating into rate limits and GPU
capacity. None of this shows up on a standard IT dashboard unless someone has
specifically built the instrumentation to see it.
And the service levels question, always the most important
one, has become the most morally loaded. When an AI agent makes a decision that
affects a customer, about a loan, a
medical triage, a service entitlement, a pricing offer, that customer has a right to understand it,
challenge it, and have a human correct it. This is not only a regulatory
requirement in an increasing number of jurisdictions. It is the operational
definition of customer focus in a world where the agent, not the employee, is
the primary interface. A Customer Service Commitment in the AI era must include
commitments about explainability, human override, and recoverability, not just
turnaround time and accuracy.
The Measurement Culture the AI Era Demands
Bringing this together, what does operational excellence
actually look like for an organisation running AI at scale?
It looks like flow management that tracks not just whether
transactions were processed, but whether the AI agents that touched those
transactions acted within defined parameters, and that surfaces exceptions
rather than silently absorbing them.
It looks like capacity management that instruments AI
resource consumption with the same rigour that a hazardous chemical plant
instruments its pressures and temperatures, understanding not just current utilisation,
but growth trajectories, shared resource constraints, and the hidden consumers
that creep up over time.
It looks like Customer Service Commitments that extend into
the AI layer, that define not just what
will be delivered, but how decisions will be explained, how errors will be
corrected, and how human accountability will be maintained even where AI is the
primary actor.
And it looks like an organisation where data is used not to
satisfy bosses or produce compliance artifacts, but as a genuine tool for
continuous improvement by everyone at every level. Where a shift in a trend
line is treated as a signal worth investigating, not as noise to be explained
away. Where prevention is valued as much as heroism. Where the excitement of
building is matched, at last, by the discipline of running.
The Hardest Part Has Not Changed
In the end, the measurement framework, however well designed,
is only as good as the culture that uses it. And culture is stubbornly human.
The data is the easy part. The hard part is persuading organisations and the
people within them to use data as a tool for honest self-improvement rather
than as a performance to be staged for external audiences.
That challenge has not changed in sixteen years. It will not
change in the next sixteen either. What changes is the cost of getting it
wrong.
Give the people the facts, about their processes, their
agents, their customers, their capacity, their failures, and their potential, and they will, if the culture is right, do the
right thing.
That is still the bet. It is a harder bet to lose than it
has ever been. But it is the only bet worth making.
"The customer does not care about your dashboard. They care about what happened to them. Those are not always the same thing."
Retaled Posts



