AI at Work Isn't a Replacement. It's a Second Pair of Hands.
Two questions, one answer
Two questions come up whenever a company starts talking seriously about AI. The first comes from whoever runs the business: "Is this actually ready, or are we still in demo territory?" The second comes from someone on the team, usually after the meeting, usually quieter: "So… am I being replaced?"
They deserve the same honest answer. AI at work, done properly, is a second pair of hands. It takes the part of the job that was never the point (the typing, the first draft, the lookup) and leaves the judgment where it was. That's the version worth building, and it's what this article is about. Where it goes wrong, and it can, is covered too.

What changed, and why "now"
For a long time, AI in business meant a slide deck. A vendor would show a chatbot answering three rehearsed questions, everyone would nod, and nothing would ship. Three things are different now.
The models got good enough, most of the time, at the boring middle of knowledge work: drafting, summarising, finding the relevant paragraph in a 40-page document, writing the boilerplate half of a function. Not the hard part. The part everyone does and nobody enjoys.
They got cheap enough to run on routine volume, not just on a pilot that somebody has to babysit.
And the plumbing matured. A model can now call your order system, read your actual policy document, and hand a conversation to a person with the context attached. That's the difference between a demo and a thing that runs on a Tuesday at 2 a.m.
None of that is the reason to start now, though. The reason is that the hard part of adopting this isn't technical. It's a team working out, over months, which tasks to hand over, which to keep, and how to check the output. Waiting for a better model doesn't shorten that. It just delays the start.
What a second pair of hands looks like in real jobs
Here's what people who do it every day say, each one linked, so you can read the whole thing and judge for yourself.
A developer. Simon Willison, who has written software professionally for over 25 years, describes an LLM as an over-confident pair programmer: fast at looking things up, happy to do the tedious bits, and sometimes wrong in ways that are subtle. His rule is that he runs and tests everything it produces. The tool doesn't replace his experience; his experience is what makes the tool useful.
An engineer on support duty. Engineers at Aha! take a week on customer support roughly every seven weeks. Winfred Wolfgram now starts each ticket by describing it to an AI assistant, which decides which logs to pull and which queries to run, and he runs them. He built five small debugging tools for it to use, and he strips customer data out of the ticket before anything goes in. He says he's happier on support and spends less time on minutiae. Note who's in charge of the arrangement, though: he built the tools, he decides what the model sees, and he's the one at the keyboard.
A doctor. Dr. Eric Boose, a family physician at Cleveland Clinic, has used an AI note-taker for about two years. As he told KFF Health News, in a piece republished by the Cobb County Courier, it listens during the visit and drafts the summary, which means he sits and listens to the patient instead of typing, and he gets home earlier. The same report is equally clear that these tools can miss things or invent them, and that clinicians have to review and edit what comes out.
An accountant. Brian Davis, a CPA who runs his own firm in Florida, uses AI for tax research and for building client deliverables from his own templates. The Journal of Accountancy reports that a complex M&A question took him about two hours, far less than it used to, and that he makes the final decisions himself before anything reaches a client.
A teacher. Sandy Mangarella has taught high-school English for 44 years. She uses AI a few times a week to brainstorm lessons and get a second opinion on essays. Recommendation letters that used to take up to an hour now take minutes, and in her view they're better than before.
A translator. Tom Gally translates Japanese to English professionally. His workflow runs the text through several models, picks the sentences he likes, then checks every paragraph against the original himself. He says the workflow is aimed at quality, not speed, and that it only works because he reads the source language and can tell when the machine is wrong.
Six people, one shape. The tedious part goes to the model. The checking and the responsibility stay with the person. None of them describes handing over the judgment, and the two who talk about it most, Willison and Gally, are blunt that the whole thing depends on knowing enough to catch the model being wrong.

The numbers are more mixed than the stories
A study of more than 5,000 customer-service agents followed what happened after an AI assistant was rolled out. Issues resolved per hour went up about 14% on average. The gain was concentrated in the newest agents, who improved by roughly a third, while the most experienced barely moved. The authors suggest the tool passed on what the more able agents already knew to the newer ones.
A field experiment across 66 companies found that workers who used an AI assistant spent about two hours less on email every week, and worked less outside normal hours. Beyond that, what they worked on didn't measurably change.
Two years after ChatGPT launched, Danish administrative records show no measurable effect on workers' hours or earnings. New tasks did appear: generating content, overseeing the AI, integrating it.
Then there's the study at BCG with 758 consultants. On tasks the model was good at, people with AI were around 25% faster and their work was rated higher. On a task deliberately chosen to sit outside what the model could do well, the AI users got it wrong more often. Correct answers fell from 84.5% to 70.6%, because they trusted the model's analysis. More trusting, more wrong. The tool is uneven, and the person has to know where the edge is.
So: real gains, mostly for the less experienced, on the tasks the model is good at. No sign so far of people losing hours or pay. And a clear way to lose, which is trusting it where it's weak.
What we see on our own projects
We build this kind of system for a living, and the pattern holds there too.
On a support agent for an e-commerce client, 68% of conversations are resolved without an operator: order status, returns, that kind of thing. When the model's confidence is low, the conversation goes to a person. Response times dropped 60% and CSAT is 4.5 out of 5. Two-thirds of the queue no longer needs a human, and that's worth saying plainly: what a company does with that freed capacity is a decision, and it's the owner's, not the model's.
On a fraud-scoring system for a fintech, the model scores transactions in under 200 milliseconds, false positives are down 40%, and 22% more fraud is caught. Nobody approves each decision at that speed. What makes it workable is that every decision is explainable and can be reconstructed after the fact, because a compliance reviewer has to be able to understand it.
On a computer-vision system across 150+ stores, on-shelf availability went up 18% and audit time fell 70%. The alerts land on the operations team's dashboard; a person still walks to the shelf.
More of this kind of work is in our portfolio.

If you run the team: how to start without breaking anything
- Pick one slice. One boring, high-volume, low-stakes task. "Where's my order," first-line calls, the weekly report nobody reads until it's late. Not the whole operation.
- Keep the person where a mistake is expensive. Money, declined customers, anything that can't be undone: either a person approves it, or every decision is recorded so a person can audit it and reverse it.
- Decide what "worked" means before you start. A number and a threshold, agreed with whoever owns the budget, measured against a group that kept doing it the old way.
- Tell the team the actual goal. If it's "the same people, fewer tedious hours," say that. If it's fewer hires next year, or a higher quota with the same people, say that too, because people will work it out anyway. Rest of World's reporting from Philippine call centres shows what it looks like when the tool arrives as a monitor instead of a helper: more calls per shift, a ceiling on handling time, demerits for tone, pitch and pauses, and an agent who says he thinks of the AI as his boss.
If you're on the team: what to do this month
Start with the task you dread. The status update, the first draft, the formatting. Hand it over and check the result the way you'd check a new colleague's work. Before you paste anything in, find out what your company allows. Wolfgram strips customer data out of every ticket before the model sees it, and that habit is worth copying on day one.
Keep your judgment sharp, because that's what the tool runs on. Willison tests everything it writes. Gally checks every paragraph against the source. The value in both cases comes from a person who knows enough to catch the model being wrong, and that's a skill you build by doing the work, not by delegating it.
Which is the one caution for anyone early in their career. Mitchell Hashimoto, a developer who leans on AI agents heavily in his own work, worries openly about skill formation in juniors. That isn't at odds with the support-agent study, where the newest agents gained the most: getting faster at a task and learning the craft are different things. Use the tool to go faster at things you already understand, and learn the rest the slow way first.
And tell your manager you're using it. The useful version of this happens in the open, with agreed rules about what data goes in.
Where this stops being right
- Experienced people on complex, familiar work. A randomised study by METR had 16 experienced open-source developers work on projects they knew well, with and without early-2025 AI tools. With AI they took 19% longer, while believing afterwards they'd been 20% faster. Whatever your senior people say about it, measure.
- Outside the model's competence. The BCG study above: more trusting, more wrong. If nobody on the team can check the output, don't automate that step.
- When the saved time becomes a quota. The call-centre story again. If the tool's main effect is that the same people are expected to handle more, with their tone scored, they'll treat it as surveillance, because it is.
- When it takes the part of the job that was the point. A GP who was an enthusiastic early adopter stopped using an ambient scribe after 18 months: the notes were accurate, but they no longer sounded like him, and he felt like a passive observer in his own consultations. Relief from admin is real. So is the cost when it's the wrong admin.
- When there's nothing to ground it in. A model with no access to your data, your systems or your policies is a demo. It will be confidently generic.
FAQ
Will this replace my team? Not on the evidence so far. The Danish records show no effect on hours or earnings two years in; the field experiment found saved email time and no change in what people worked on. What the evidence does show is capacity: the same people getting through more. Whether that becomes fewer hires or better work is a management decision, not a property of the tool.
Where should we start? One narrow, high-volume, low-stakes slice, with a person on the exceptions and a metric agreed before you begin. Expand only once that number holds.
I'm on the team. Should I say I'm using it? Yes. Ask what data you're allowed to put in, and keep checking the output. This works best when the whole team is doing it in the open.
The practical version of all this is short. Give the model the volume. Keep a person on the exceptions and on anything that can't be undone. Measure, because self-reports go wrong in both directions. And say out loud what you'll do with the time it frees. That's what we build at ISTRALLEN: agents and models grounded in your real data, with the person kept exactly where the judgment matters; see what we do.