Insights · Practice
Nobody wrote it down
AI is absorbing the legal work the profession wrote down. Most of what it protects as judgement was never uncodifiable, just unwritten, and that turns the automation curve from a forecast into a choice.
A designed preview of a piece still in draft. The diagrams are generated previews rather than final exports, and the article stays excluded from search indexing until it is cleared for publication.
There is a question I put to lawyers in almost every training session I run: if one piece of your working week could disappear tomorrow, which one goes? I have yet to hear anyone say drafting. They say the chasing. The follow-ups. Going into a third system to find out whether something actually happened. The answer is consistent enough that it has stopped surprising me, and the profession has built almost nothing in response to it.
That question carries more weight this year, because AI can now do, rather than assist with, a very large amount of legal work. The research on what that means has so far been about juniors. Stanford's Canaries paper found employment of 22 to 25 year olds in the most AI-exposed occupations running 19% below where it would be had it kept pace, with the effect operating through hiring rather than redundancies, paralegals sitting in the most exposed quintile and lawyers in the one below it. The piece everyone will write off the back of that is “AI is breaking the training ladder”, and the paper broadly says so itself.
The honest observation goes considerably further than juniors, because seniors do a great deal of work they have no business doing. Partners chase. Senior associates build chronologies. Everyone checks whether the thing got filed. So this piece asks the questions a practising lawyer of any seniority can act on:
- Where does this actually land in your day?
- Where can you be genuinely supercharged?
- What should you stop holding on to and hand over?
- What is in the way of doing either?
- And if you delegate, how do you still learn it? If you supercharge, how do you stop the skill decaying underneath you?
What the day is actually made of
Task inventories in this market usually arrive as a ladder, junior work at the bottom, judgement at the top. The honest version is six bands, ordered by how much of each genuinely resists being written down, and deliberately not ordered by seniority, because seniors do plenty of band-one work and pretending otherwise makes the inventory useless.
Finding and reading.Disclosure and e-discovery review. DD document review. Docket review. Needle-in-haystack searching. Establishing whether a contract exists at all, and whether you are looking at the latest version. This no longer needs humans. It does need a sufficiently sophisticated system, with checking loops and direction about what to look for, rather than a rigid list of tightly defined things to extract. That distinction explains why so many firms concluded “we tried this and it didn't work”: they tried the previous generation of extraction tooling, which needed every question specified in advance. The failure was real and it belonged to the tool generation, not the proposition. I have run this experiment myself; the due diligence risk identification work we published sits exactly here.
Producing from a known form. Drafting from precedents. Form completion. Note-taking, summarising, reporting, building timelines from a document set. Already substantially moved, and the profession has largely accepted it.
Applying a known standard to someone else's material. Contract playbook markup. Copy clearance, reading adverts against ASA requirements and feeding back the changes needed, which is regulatory playbook markup and behaves like the contract version. Compliance checking. Exposed by definition: a playbook is codified knowledge with a name on it.
Running the matter. Chasing people. Following up. Checking whether things have actually happened. Jumping between tools. Project management, filing and administration, Companies House and Land Registry. Exposed, wanted, and ignored, and it deserves more attention than the rest of the inventory put together.
Analysis.Quantum assessment, heads of loss, whether the numbers line up. Finding the gaps in the other side's argument. Benchmarking and “is this market”. Contested, and moving, because much of this is data analysis wearing a wig.
Judgement and the human-facing work. Triaging what matters in a matter. Risk assessment. Case strategy. Advising. Interviewing witnesses. Negotiation. The residue sits here, and it is smaller than the profession thinks.
Unwritten, not unwritable
The Canaries paper explains its findings through codified versus tacit knowledge: AI substitutes for formal, documented, teachable knowledge and complements experiential knowledge. That framing is useful and incomplete, because it treats the codified/tacit split as a fact about the work. Mostly it is a fact about what the profession has bothered to write down.
Most of what lawyers call tacit knowledge is uncodified rather than uncodifiable. Nobody has sat down and written out how you triage what matters in a matter, how you decide which threads are worth following, what actually drives a risk call. It could be done. It would take real effort, and nobody has ever been paid to do it. So it stays in people's heads, gets transmitted by osmosis, and gets described as judgement.
Two things follow from that reframe. The training problem stops being an AI problem: we never codified the thing juniors were supposed to be learning, so the only transmission mechanism was doing the work, and taking the work away exposes the fallback we never built. And the automation curve turns from a forecast into a choice. What is exposed is largely what has been written down, so a firm that codifies more will find more of its work moves, on its own terms and to its own benefit, rather than waiting to discover which work a vendor moves for it.
The top of the chain, split honestly
The judgement band is usually treated as one protected zone. Split it honestly and it comes apart.
Triage, risk assessment and case strategy are uncodified, not uncodifiable. Experience means you can tell what happened before, and that experience could be captured if anyone collected the data, followed matters up and recorded outcomes. Nobody does. The obstacle is a missing feedback loop, not a mystery.
Interviewing and witness work is human-facing, but its output is codified. Taking a statement involves a person in a room; drafting one does not.
Negotiation genuinely resists, and the reason is more precise than “it's tacit”. A negotiation playbook is entirely codifiable. The constraint is that codifying and disclosing are different things, and what protects negotiation is adversarial information asymmetry. Nobody opens with their ideal position, because stating it gives it away. You go in with your standard, knowing what you will accept, what you will not, and what you are prepared to trade, and you avoid signalling that you care too much about any one point, because a competent counterparty will use that to extract something else. All of that could be written down completely. What cannot be automated is the live reading of the other side, and how much of that comes down to knowing your particular client and what they will actually live with.
So the honest axis at the top of the chain runs from codifiable-but-uncodified, which is fixable, and the fix is a data problem, to adversarial and relational, which writing more down does not touch.
The work nobody defends
The bands order by resistance to codification, which is the right ordering for the argument and the wrong ordering for demand. Go back to the training-room question: the work lawyers most want taken away is running the matter, and it gets a fraction of the attention paid to drafting and markup. Four reasons why.
It is misclassified as low-value work, which is a category error about status rather than a judgement about cost. Running the matter is not legal reasoning, so it gets filed as administration and treated as beneath discussion. But time is time. The profession measures work by intellectual prestige and then wonders where the day went.
Nobody defends it, which makes it the least resisted automation available. Propose automating drafting or markup and you get an argument: AI cannot really draft, you have not seen what it does to a warranty schedule. Propose taking away the chasing and nobody mounts a defence, because nobody believes chasing is a professional skill. So the band with the least cultural resistance and the highest expressed demand is the one firms deprioritise, which is close to exactly backwards for anyone sequencing an AI programme.
It is not document AI, so the vendors and the commentary both miss it. Running the matter needs systems that reach across tools, hold state and act, rather than models that read documents well. The market is document-centric because the products are, and the discourse follows the products. This is agentic and integration work, and it is the least covered part of the market relative to how much of the day it eats.
And the commercial argument is stronger here than anywhere else in the chain. The persistent ROI problem in legal AI is that time saved on drafting does not convert to money under the billable hour, which is why every vendor ROI study ends in an argument about conversion assumptions. This band does not have that problem, because most of the work was never billable in the first place. It is written off, absorbed, or done at nine in the evening. Returning it is a straight gain with no billing paradox and no cannibalised revenue.
If you need one argument that survives contact with a finance director, use this one.
The numbers behind the demand are stark. BigHand's 2026 workflow report found 67% of firms still delegating work by direct email in whole or part, with only about a fifth using a structured workflow system as the primary route, and the creation of new junior administrative roles collapsing from 63% of firms in 2022 to 28% in 2026. BigHand sells the workflow product every one of those findings points toward, so discount the framing, but the delegation and workforce data stands, and the last figure matters twice over: the people who used to absorb this work are no longer being hired, and nothing has replaced them except lawyers doing it themselves.
Task switching is not a task
One item in that band behaves differently from the rest and is worth separating. Task switching and tool switching are a tax on every other item of work rather than items of work themselves. They are what makes a six-hour day produce three hours of output, and they never appear in any task inventory because they are the gaps between the tasks.
The reason a lawyer jumps between five systems to establish one fact is that the information is not in one place. That is a context problem, and this band is what the missing context layer costs, measured in hours out of a working day rather than in architecture.
One chain, two tracks
One split does need making, and it is contentious against transactional rather than practice area against practice area. Transactional runs across real estate, corporate, commercial and finance; contentious runs across disputes, arbitration, claims and arguing a case. These are genuinely separate skills, and which one you go into changes what your years in practice teach you. The research treats “junior lawyer” as a single category, and it is not one.
Two things fall out of laying the bands against the split.
The tracks converge at the bottom and diverge at the top. The finding-and-reading band is nearly identical on both sides: disclosure and docket review on one, DD review and contract-hunting on the other, the same underlying task of finding the thing in the pile. By the judgement band they have nothing in common: transactional judgement is dominated by negotiation, contentious judgement by case strategy, settlement calls and witness handling. Automation lands on both tracks in the same place, and what survives is precisely what makes a corporate lawyer and a litigator different people.
And the usual safety assumption runs backwards. Transactional work has been travelling the codification road for decades: precedents, standard forms, playbooks and market-standard positions are all codified artefacts the profession built deliberately, which is exactly why so much transactional work has already moved. Contentious work has far fewer of those artefacts, which makes it look safer. It is further behind on codification, which means the headroom is larger: quantum analysis, gaps in an argument, chronology building and claims checking are all codifiable and mostly uncodified.
| Band | Transactional | Contentious |
|---|---|---|
| Finding and reading | DD review; does the contract exist; is it the latest version | Disclosure and e-discovery; docket review; needle-in-haystack |
| Producing from a form | Precedent drafting, form completion. Furthest along | Chronologies, timelines, witness statements, submissions |
| Applying a standard | Playbook markup, copy clearance. Playbook-native | Pleadings against procedural rules; claims against a standard |
| Running the matter | CPs, signing, completion, Companies House, Land Registry | Court deadlines, directions, disclosure windows, bundles. Hard-coded by rules, so arguably more exposed |
| Analysis | Is this market; benchmarking a term | Quantum, heads of loss, gaps in the argument. Much richer |
| Judgement | Negotiation dominant | Case strategy, settlement calls, witness handling |
The perforated ladder
There is a genuine objection to handing over the reading, and it deserves meeting head-on: reading a great many documents and getting a feel for what is going on is valuable, and following threads to work something out probably does leave you understanding the matter better. Preserving inefficient work as a training mechanism is still the wrong response. If the only way a junior learns to read a matter is by grinding through the documents, that was never a training programme. It was a by-product we got lucky with for a century.
Draw the six bands as ladder rungs and mark what goes. The rungs that go are scattered, because the bands order by resistance to codification and not by seniority: finding-and-reading and producing-from-a-form go, applying-a-standard holds for now, running-the-matter goes, analysis and judgement hold.
The ladder did not get shorter. It got holes in it.
That is why “just start juniors higher up” is not available as an answer: what remains does not join up. Each rung taught something specific. What a document set feels like and which threads are worth pulling. What good looks like and the shape of the instrument. Why the standard says what it says and where it bends. How a matter actually moves and who does what when. How to test a number and find the gap in an argument. What this client will live with. Remove scattered rungs and the gaps in the teaching are scattered too.
One thing worth testing rather than asserting: is the feel you get from volume reading a real artefact of the volume, or nostalgia? I do not know, and I have not seen evidence either way. It is exactly the kind of question a firm could answer with its own data, if it collected any.
A skill you never had, and a skill you did
The published research raises the over-reliance question and leaves it lying there. It splits into two failure modes pointing in opposite directions, and conflating them is why most commentary on it is useless.
Delegating risks a skill never forming. You hand the band to AI, you never do the work, so you never acquire the thing doing it used to give you. This is the trainee problem, and it is about formation.
Supercharging risks a skill decaying and judgement going quiet. You already have the skill, but the tool is usually right, so you gradually stop checking, stop forming your own view first, and stop noticing when it is wrong. This is the senior problem, it applies to people who were never at risk from the first failure mode, and it is about maintenance and automation bias.
There is an obvious answer here that this piece cannot give: keep doing the work by hand so you learn it. That is the by-product training model I have just argued against, and if the answer is “do it manually”, the piece eats itself. The answer that holds is that the learning has to be designed in deliberately, and the evidence says the outcome depends on how you use the tool, not on whether you do the work yourself.
The evidence is unusually good. Shen and Tamkin at Anthropic ran a preregistered trial in which developers using AI to complete a task on an unfamiliar library scored 4.15 points lower out of 27 on a comprehension test afterwards, with no speed benefit. The part you can use is behavioural: participants who generated an answer and then asked a follow-up question to understand it learned the material, and participants who generated an answer and stopped did not. Same tool, same output, and the difference was one additional question.
The Minnesota randomised trial (Bednar, Cleveland, Erbsen and Schwarcz) found AI at the front of a legal task, on the user's own terms, lifted synthesis quality by 59.8%. The same participants applying AI at the end of the session, under time pressure, made strong work worse: weak memos improved and the strongest degraded. Same tool, opposite outcomes, decided by sequencing.
Microsoft's scaffolding field experiment ran the lesson at organisational scale. Imposing a rigid step-by-step AI protocol on 388 Gap Inc employees produced significantly worse documents and made producing no document at all eight times more likely, while training people to treat the tool as a thought partner roughly doubled the odds of a perfect score. Telling people the steps makes it worse; teaching them the stance makes it better.
The people this is happening to can already see it. Eight out of eight groups at a Harvard workshop of general counsel raised skills atrophy unprompted, whatever topic they had been assigned. Thomson Reuters' survey of 1,874 law students found 74% fear over-reliance eroding their core skills, and they rate AI's effect on independent judgement at net minus 31, the worst score on the card.
Three moves follow, and none of them is “use it less”.
Form your own view first, then compare. The sequencing result and the follow-up-question result are the same finding twice: engage before or after the output arrives and you learn, defer and you do not. The cost is minutes. The difference is whether you learned anything.
Treat verification as the skill and teach it as one. If juniors are checking AI output rather than producing it, checking is the competence, and nobody currently teaches it. The BigHand data shows where the bill lands otherwise: 46% of firms say AI output needs additional supervision and 45% report more time spent checking, which is the time saved reappearing as supervision by the most expensive person in the building.
Codify what you learn as you go. This is where the individual answer and the institutional answer become the same act. Write down how you triaged the matter and why the risk call went the way it did, and you have protected your own reasoning and built a piece of the thing the firm never built. Anthropic's own usage research points the same way: in agentic work, people still make about 70% of the planning decisions, and domain expertise is what makes the tool productive. The expertise survives. The question is whether it gets written down.
Leverage without the pyramid
For any band delegated wholly to AI, the leverage changes in kind rather than degree, because there is no people cost at all. The traditional model earns leverage by putting cheap humans under expensive ones and marking up the difference. Move a band entirely to AI and the return on that band goes somewhere the partnership model has no way to price, and a firm that never built the pyramid has a structural cost advantage on exactly that work.
Bommarito frames the apprenticeship problem as a financing failure in The Last Partnership; PwC models the pyramid gaining an agentic layer beneath every human tier. Both arguments run at firm level. The band inventory says which work it happens to first.
The office, briefly
There is a rival explanation for the entry-level decline worth naming. Lambert and Schindler argue it is driven by remote work rather than AI, and both their paper and the Canaries paper classify legal work as teleworkable. My instinct is that law is not a high-WFH profession in practice, which means the disagreement is really about classification against actual behaviour, and that is worth testing rather than assuming.
Either way, the useful move is to turn it into the cultural question every reader is currently arguing about. Firms are pushing people back into the office because they believe it is better. Better for what? If the answer is training, then presence only trains anyone if someone intentionally trains the trainees while they are there. Bums on seats is not a training programme either.
A decision you can reverse
If most of what looks tacit is merely uncodified, and if experience could be captured by collecting outcomes and following matters up, then the constraint on legal AI stops being model capability. The constraint is that nobody has built the place where that knowledge would live: the record of how matters were triaged, which threads got pulled, what the risk call was and how it turned out.
The work is uncodified because nobody was ever paid to write it down. That is a decision, and it is one a firm can reverse, band by band, starting with the work its lawyers would thank it for taking away.
Source notes
Employment and mechanism findings from Brynjolfsson, Chandar and Chen, Canaries in the Coal Mine (Stanford Digital Economy Lab, August 2026 revision). Delegation, junior-role and verification figures from the BigHand Legal Workflow Leadership Report 2026 (800+ senior leaders; vendor-produced, so the framing is discounted and the structured-delegation figure is quoted as “about a fifth” because the report's chart and narrative disagree). Skill formation: Shen and Tamkin, How AI Impacts Skill Formation (preregistered RCT, n=52; the usage-pattern split is reported as direction only because the underlying clusters are small). Sequencing and revision effects: Bednar, Cleveland, Erbsen and Schwarcz, University of Minnesota working paper, April 2026 (preregistered RCT, n=91). Scaffolding: Farach, Cambon, Tankelevitch, Hsueh and Janssen, Microsoft field experiment, March 2026 (n=388). Practitioner and student sentiment: Harvard CLP, Moving from Adoption to Responsible Impact and the Thomson Reuters Institute 2026 Law Student Pulse Survey (n=1,874). Planning-versus-execution split: Anthropic, Agentic Coding and Persistent Returns to Expertise (June 2026). Firm-economics framing: Bommarito, The Last Partnership, and PwC, The New Rules of Legal Services (2026). The remote-work rival explanation is Lambert and Schindler (2026). Diagrams are generated previews of the working figures.
Continue reading