The Bluestem Mutual water tower came up before the exit signs did. I was still in the left lane off the interstate when BLUESTEM MUTUAL appeared above the tree line, blue letters faded to the color of old denim on a white tank that hadn’t been repainted since the nineties. I’d worked this exit before, not for Bluestem, different client, different decade, but the geometry was the same: two-lane county road, a stand of elms that needed trimming, and at the end of it a campus that looked smaller than the revenue it produced.
Bluestem writes about nine hundred million in premium a year. Workers’ comp, commercial property, general liability, mid-market accounts across fourteen states. The comp is thirteen of the fourteen, everywhere but home: Ohio keeps workers’ comp for itself, state fund only, and has since before Bluestem existed, so the company sells everything else on its own doorstep and comp one state line over. Three buildings on forty acres, parking lot that wanted repaving, a water tower with its name on it. The tower is real infrastructure, not signage. When Bluestem built the campus in the fifties, the county’s water pressure couldn’t feed a sprinkler system to the standard Bluestem’s own underwriters demanded of anyone else, so the company built its own supply rather than carry a risk it wouldn’t have accepted from a customer. Somebody in facilities still has it inspected every year. I put the campus at maybe twelve million in replacement cost. I’ve been doing this long enough that my brain prices things before I’ve decided whether I care about them. My wife called it a professional deformation once, over dinner, and then we talked about something else.
I’d been doing independent consulting for four years by then, mostly insurance and financial services, mostly the AI projects that carriers kept announcing and then couldn’t scope. The pitch was always roughly the same: we need to modernize, we have a lot of data, we’ve heard good things about large language models. My job is to figure out whether any of that is true before someone spends three million finding out the hard way.
Priya Raman had called three weeks prior, on a Thursday, while I was finishing a six-month engagement at a regional carrier in Cincinnati. She said Bluestem needed an AI architect for claims ops and she needed someone who wouldn’t oversell them. I told her my rate was twelve thousand a week. She said fine without pausing. That either meant the problem was bad enough that the price had stopped mattering or the budget was already approved, and I figured I’d find out which.
I parked in visitor and sat in the car for a minute. When a VP of claims operations calls an outside architect, she’s been telling her own IT team what she needs for a while and not getting it, or she’s been told something from the executive floor and doesn’t trust her team to tell her whether it’s real. Maybe both. Priya had the phone manner of someone who had already done the internal cycle and was now paying outside rates to short-circuit it.
The lobby had burgundy carpet, dark wood furniture in a pattern that had given up trying, and a receptionist who called Priya before I’d finished spelling my last name. I put a renovation at maybe two hundred thousand if they ever got around to it. They hadn’t.
Priya was there before I’d found a seat. Forty-four, dark hair pulled back, a pace that meant she had already been three places before this meeting and was noting two more she needed to get to after. We shook hands and she walked me to a conference room without small talk, which I appreciated. Marcus Webb, the IT director, was already there, seated, laptop open, wearing the expression of a man who has stopped hoping for a quiet Tuesday and is now just managing for the absence of a catastrophic one. He runs twelve systems with thirty people, which sounds manageable until you learn that six of the twelve talk to each other through scheduled file drops and one of them runs on a server nobody wants to migrate because the last migration cost four days of claims payments. I didn’t know that yet. I found it out later.
“We need a chatbot for claims,” Priya said. She said it like a sentence she’d been carrying for a while and was glad to put down.
I opened my notebook. “What problem does it solve?”
“Cycle time. We’re averaging eleven days from assignment to a coverage determination on a complex claim. I’ve been asked to bring that down.”
“By whom, and to what target?”
“Executive team. The target is still being set.” She moved on before I could follow up. “Marcus.”
Marcus turned his laptop to face me. A ticket queue, color-coded by age. “My team handles about four hundred support calls a month from adjusters who can’t find information across the systems. Better tooling, some of that deflects.”
“How many of the four hundred.”
“Maybe half. I don’t have a harder number than that.”
I wrote two hundred, question mark.
The phone on the conference table rang. Priya answered and put it on speaker. A voice came through, slightly too far from his microphone. He was glad I was there. This was an exciting initiative. He used the word innovation twice. He wanted to make sure we weren’t just thinking about internal efficiencies, that the customer experience side was also on our radar. Then he said he had another call and the line went quiet.
Nobody said anything for a moment.
“SVP of Strategy,” Priya said.
“Does he have a metric he’s watching.”
“He has a perspective,” she said. That was the most useful thing anyone said all morning.
Three requests: cycle time, ticket deflection, innovation. I’ve sat in this meeting, or close enough to it, at a health plan in Phoenix, a specialty lines shop in Hartford, a workers’ comp carrier outside Minneapolis. The solution, usually a chatbot or an AI assistant, gets chosen before the problem is named in a way you can measure. The SVP’s innovation would dissolve at first contact with a budget review. That left two requests with numbers attached, even if the numbers were soft.
Eleven days was a number Priya knew without looking at anything. That meant someone was counting, and the count was going the wrong direction. Four hundred tickets a month was a number you count because it’s consuming something. I didn’t know yet what was driving eleven days, or whether any of those tickets were fixable without rebuilding something structural. Not enough to commit to anything.
“Two weeks,” I said. “Discovery only. I want to understand where the time actually goes in a complex claim before we talk about what to build.”
Priya nodded. “Start with Carol Jansky. Senior adjuster, second floor. She’s been here longer than most of our systems. If she thinks something is worth doing, the floor follows.”
“And if she doesn’t.”
“Then we’ll have learned something,” Priya said. I think she meant it both ways.
Two weeks. Eleven adjusters, two supervisors, one IT admin who managed DocVault’s permissions structure and had strong opinions about the adjusters who requested access without knowing what they needed it for. I walked the second floor twice: once on a Tuesday when the pace was normal and once on a Thursday when three claims had hit the same escalation threshold and the supervisor was on her fourth phone call before nine-thirty.
The adjusters’ complaints ran together: too many systems, not enough time, please don’t give us another login to remember. I heard about ClaimCore from six different people without asking. I started asking on purpose, to see what the narrative settled into, and it always settled in the same place: the migration went live in 2019, data was lost or corrupted in ways that didn’t surface for months, and adjusters spent the following year doing cleanup work that wasn’t supposed to exist. Three people mentioned it in the same breath as “another IT project.” One mentioned it in the same breath as my name, which I noted.
Carol Jansky sat in the northeast corner of the second floor: two monitors, a paper Rolodex with actual paper cards sorted by hand, a yellow legal pad with blue ink, a framed photo I didn’t examine. Thirty-three years at Bluestem, which meant she had outlasted four system migrations, two rebrands, and whatever came before ClaimCore. She looked at me the way you look at someone who has shown up with good intentions and a clipboard and a track record you don’t know yet.
“I was here for ClaimCore,” she said before I’d fully sat down.
“How was it.”
“We went live in October 2019. I spent most of 2020 reconciling claim histories the migration had dropped or garbled. Endorsement attachments that simply weren’t there anymore. The vendor called it a successful go-live.” She looked at me steadily. “So I’ve developed some opinions about IT projects.”
I told her I wanted to watch her work through a complex claim and that I’d stay out of the way. She said she had one open right now, general liability, she’d narrate it if I didn’t mind. I said thank you and took out my phone and started the timer before she turned back to her screen. I felt mildly bad about that.
The insured was a restaurant group, four locations in Indiana. The claimant wasn’t theirs: a driver for their produce subcontractor, laceration, right hand, caught on prep equipment while wheeling a delivery through the kitchen. Reported three weeks prior. She pulled it up in ClaimCore, scanned the intake notes, then switched screens to the policy. “General liability, and the account has a kitchen-operations exclusion negotiated onto it,” she said. “I need to know whether that exclusion reaches this injury, because if it does we’re writing a denial with an appeals path attached, and if it doesn’t we owe a defense, and getting that wrong is expensive in both directions.”
She opened DocVault. The policy was 178 pages. She didn’t start at page one. She pulled the declarations page and the schedule of forms first, the list of every endorsement attached to the policy, and ran her cursor down it until she found the kitchen-operations exclusion, manuscript language, written for this account. Nobody reads the base form, she told me without looking up; the base form is the same boilerplate on ten thousand policies, and the endorsement schedule and whatever got negotiated onto the account are where the ninety minutes go. The endorsement was stored as a separate document, which meant a second search. “This is the part where I usually get a coffee,” she said.
The endorsement ran fourteen pages. She ran Ctrl-F on service inside it: twenty-six hits. She read through the relevant paragraphs, made a note on the legal pad, then went back to the definitions section and read it twice. “The exclusion applies to injuries arising out of food service operations,” she said. “Whether prep work counts as service is the question. The definitions section is doing a lot of work here and not doing it clearly.”
She opened a legal research tab and found a state appellate decision, Indiana, that had construed the same phrase in similar manuscript language. She read without speaking for a while.
I watched my phone. Sixty-seven minutes.
“Prep isn’t service under the endorsement’s own definitions, and the Indiana court read the phrase the same way,” she said. “Exclusion doesn’t apply. We owe the defense.” She closed two tabs and wrote two lines on the pad. “That’s it.”
“How long did that take you.”
She looked at her watch. “About an hour and a half. That’s normal for this type.”
My phone said eighty-eight minutes. Near enough.
“Every complex claim goes through that.”
“Every one with endorsements or anything state-specific, which is the majority of the ones that require actual judgment.” She set down her pen. “I have eleven of those waiting on coverage review right now. That’s my week, before I do anything else.”
Ninety minutes times eleven was sixteen and a half hours. One adjuster, one week, all of it spent reading PDFs to figure out what was covered before she could decide anything or call anyone.
“What would help,” I said.
“Someone who already knew the answer.” She said it without heat, like a thing she had said before and gotten tired of saying. “I’ve been asking for that for four years. What I got was a new login to DocVault and a search function that returns forty results and makes me read all of them anyway.”
I wrote it down word for word.
Dale Kessler’s office was on the third floor, corner, two windows looking toward the parking lot. Clear desk, good chair, a whiteboard with three columns of figures I was too far away to read. He’d been CFO for nine years, long enough to have signed off on the ClaimCore migration, which I knew going in, and which I figured he knew I knew.
He didn’t offer coffee. He waited until I’d sat down and said, “Give me the number.”
“I don’t have one.”
“You’ve been here two weeks.”
“Doing discovery. I have a baseline, not a projection.”
He looked at me the way you look at an invoice you weren’t expecting. Then: “What’s the baseline?”
“Bluestem closes around fourteen thousand complex claims a year. Anything that requires coverage review before the adjuster can make a decision. I watched Carol Jansky work through one and timed it: ninety minutes of reading policy documents out of DocVault to make a single coverage determination. She said it’s typical for that category.” I found the page in my notebook. “Fourteen thousand claims, ninety minutes each: twenty-one thousand adjuster hours a year just for coverage review. At eighty-five dollars loaded cost per hour—”
“You’re low,” he said. “With benefits and the performance bonus structure, loaded cost is ninety-two dollars.”
I crossed out eighty-five and wrote ninety-two. “Ninety-two times twenty-one thousand is $1.93 million. Call it $1.8 million to stay conservative.”
“Call it $1.9 million.” He hadn’t touched a pen. “That’s just the reading.”
“Coverage review only. It doesn’t include rework when someone gets the call wrong, or claims that sit idle because the adjuster is mid-review on a different one. The $1.9 million is the floor.”
He tapped two fingers on the desk, once. “Everyone who comes in here with one of these projects tells me they’re going to cut costs by forty percent. Sometimes thirty. Sometimes sixty, depending on how they think I’ll react.”
“I don’t have a percentage. If you want one, I can make one up, but you’d see through it and we’d both be annoyed.”
He was quiet for a moment. I thought I’d pushed too far.
“What do you need?” he said.
“Four more weeks. The policies are complex, the endorsements have exceptions to their exceptions, and Carol’s thirty-three years of context is doing work that isn’t obvious from watching her. I need to find out whether the reading step is actually automatable or whether it only looks that way from a distance. I’m not going to tell you what to build until I know.”
“And at the end of four weeks?”
“A feasibility assessment with a confidence level. If the answer is yes and the confidence is high, you fund a build and we can talk ROI then. If the answer is uncertain or no, you’ve spent forty-eight thousand dollars finding that out instead of two million.”
He picked up a pen and set it back down. “Forty-eight thousand is your rate times four weeks.”
“Yes.”
He looked toward the window. The parking lot, the row of blue Bluestem signs along the perimeter fence. A long moment. “Carol’s been asking for better tools since I took this job,” he said.
“She mentioned that.”
“Did she mention ClaimCore?”
“First sentence.”
He pulled a folder from his top-right drawer and set it on the desk without opening it. “Approved. Discovery, not a solution. Four weeks.” He tapped the folder once. “Come back with something I can act on. Not a presentation. A decision.”
I said I would and stood. He was already opening the folder.
The SVP of Strategy arrived with a printed agenda for a meeting that had no agenda. He was tall, and he said “we need to be thoughtful” when he meant he hadn’t decided yet. I’d met him twice before that afternoon: once in the hallway during week two, once on a video call where his camera framed the ceiling more than his face. Carol called him Richard in hallways and “the SVP” in every other context, and nobody corrected either name.
We were in the second-floor conference room with the broken blind on the east window, the one that threw a wedge of afternoon sun slowly across the table. Priya sat at the head with a legal pad. Marcus had his laptop open and was building a notes document that would be meticulously organized and nearly useless for decisions. Carol sat nearest the door, which I’d learned by then was preference, not habit.
The question on the table: if we built this, how would we know it worked.
Richard opened. “My position is that any AI touching our claims process needs to be one hundred percent accurate before it goes near a claim.” He said it carefully, the way people deliver sentences they’ve practiced, and the room absorbed it without a word. In insurance, that sentence sounds like responsibility, like the thing a careful executive says to protect his organization.
I wrote it down. Then I asked what accuracy the current process achieved.
Silence. The particular quiet of four people who have never needed to put a number on the thing they already ran.
Carol broke it. “We get audited on a sample. Quarterly findings.” A pause. “I couldn’t give you an overall accuracy rate.” She said it without apology, without hedging.
Marcus looked up from his laptop. “I don’t think we’ve ever structured findings as a claim-level accuracy number.”
So we had a baseline for cost and nothing comparable for accuracy on the reading itself. Richard’s demand was serious and I wasn’t going to argue it away. But it had been measuring a proposed system against a standard we had never applied to the existing one. Naming that took six minutes.
We spent the next two hours building something better. Priya pushed every phrase that wasn’t specific enough; Marcus drafted and revised in real time; Carol kept returning to the adjuster’s actual workflow, which was the thing neither Marcus nor I had lived. At one point Marcus wrote “acceptable error rate commensurate with adjuster performance” and Priya read it back aloud, slowly, the way she read things she was about to reject. “What does that commit us to,” she asked, “specifically?” Marcus deleted it and we moved on. Richard contributed questions that were harder to answer than they were to ask. What emerged, paragraph by paragraph, was a set of criteria we could actually measure.
We would assemble a labeled sample: two hundred closed claims, manually reviewed by Carol’s team, ground truth established before any model touched the files. Carol did the math herself and got nine thousand dollars; Marcus reran it at the ninety-two-dollar loaded rate and got closer to eleven. “That’s real money,” Carol said. It was. I told her it was also the only way to have an acceptance bar that wasn’t a guess. She considered that for a moment, then put a check next to the number on her notepad.
She came back to the ninety-five percent threshold a few minutes later. “How’d you land on ninety-five?” I told her we didn’t so much arrive as land: ninety-eight sounded rigorous until we acknowledged we had no baseline to measure it against, and ninety would accept a system that made errors we couldn’t comfortably explain to an auditor. Ninety-five was the floor below which we’d struggle to justify deployment, and above which we had a reasonable argument that the tool was performing at least as well as the process we’d already been running for twenty years. She wrote it down and didn’t argue.
The other criteria followed from that framing. Every summary had to carry source citations that a human adjuster could verify by pulling the original document immediately, in under a minute. The assistant drafted. The adjuster decided. Always. That last part wasn’t a design principle subject to renegotiation during build; it went in the document as a structural given.
Richard asked about error handling. One sentence: handling would be proportionate to risk, meaning a missed policy limit on a $2 million liability claim triggered a hard stop and a human flag, and a formatting inconsistency on a $400 auto claim did not. He nodded slowly, the way that meant either satisfaction or the end of his attention. With Richard those were hard to distinguish.
Carol had gone quiet for a few minutes, running her pen back and forth along the edge of her notepad. “When it’s wrong,” she said, “how will I know?”
It’s not if it’s wrong, it’s when.
I told her we’d design for that: every citation visible, every claim scoring below the confidence threshold surfaced to the adjuster rather than passed through silently. I wrote her question down and told her it was going on the document, because it was better than anything I’d come in with. She looked at me like she wasn’t sure whether that was a compliment or a confession. Both, I told her.
Richard left at four-thirty saying he was “very encouraged by the direction.” Priya waited until his footsteps cleared the hall, then made one small mark on her legal pad. Carol was already in the elevator.
The one-page problem statement took four drafts to get to one page, which is always how it goes. The first version ran three pages and included a section on guiding principles (“prioritize adjuster experience,” “ensure regulatory compliance,” “maintain data privacy”), all of them true, none specific enough to prevent a single bad decision during build. I cut the section. Principles that don’t change what you build don’t belong in a requirements document.
What survived: the business problem in two sentences. The measured baseline: 14,000 complex claims per year, approximately ninety minutes per review, $92 per hour loaded, $1.9 million annually. The acceptance criteria from the workshop, with Carol’s question written in as a requirement rather than a comment. And the data landscape.
The document got honest about the data landscape.
DocVault held 26 million PDFs. OCR quality varied by decade in a way that mattered operationally: anything digitally originated in the past ten years was clean; 2000s-era batch scans were acceptable; late-1990s material had been processed at 72 DPI and marked complete by a platform that had done exactly what it was asked to do and nothing more. POLARIS maintained nightly replicas to a read-only environment, better than I’d expected: consistently structured, documented by someone who had clearly expected future users.
Then Marcus found the faxes.
He’d been running a file-type check across the endorsement directory in week five when he turned up a subset of older endorsement schedules (the riders that modify the base policy, often the most consequential pages in a claim file) that existed only as scanned faxes. Not PDFs built from scanned images. Thermal fax paper, physically scanned once, stored as flat image files, run through OCR software that predated most of the adjusters currently working in the building. He sent me the file count in an email with no subject line: ~340k files, commercial lines, 1998–2006. You’re going to want to look at a sample. I pulled twenty of them. The OCR on the worst ones hadn’t produced garbage output (which would have been easier to catch) but plausible output: sentences that read like insurance language and contained words that weren’t in the original document. One exclusion that read “does not include personal property” in the source image came back as “does not exclude personal property” in the extracted text. That’s a covered claim becoming a denied one, or the reverse, on the strength of a word the OCR had guessed.
I wrote that down exactly as we found it: a fact about the current data environment that any solution would have to account for before it processed a claim file that might contain one, and in commercial lines that was not a small share of claim files.
The document didn’t recommend building anything, hiring anyone, or buying software. It said: here is the problem, here is what we can measure, here is what success looks like, here is what the data actually contains. What to do about any of that was someone else’s next conversation. The word “chatbot” appeared nowhere on the page. Neither did “AI assistant” or “large language model.” Discovery was what Dale had funded, and one page of discovery was what we had.
Priya signed on a Tuesday. She read it twice, slowly, her legal pad face-down, her tell for something she thought might matter in a future meeting, and signed.
Dale signed without a meeting. The document came back through interoffice mail, the kraft paper envelope with the string closure, with a sticky note on the cover page in his handwriting: Get the fax situation scoped before build starts. No mention of the six weeks or the extension or Carol’s $11,000 labeling commitment. Just the next unresolved thing, identified and assigned.
I brought the sticky note to Priya.
“You just spent five weeks writing one page,” she said.
I told her that page was the only thing that would survive contact with engineering. The workshop notes, the interview transcripts, the three prior drafts would all compress into assumptions the moment someone opened the first development ticket; the page was what people came back to when they disagreed about what they’d meant. They always did.
She thought about that for a moment. “Dale’s going to ask about the fax situation in two weeks.”
Probably one, I told her.
I slid the signed page into a three-ring binder with a clear front pocket and drove home on Route 34, past the Bluestem water tower with its peeling paint and the town motto nobody had touched since 2019. My daughter had a swim meet Thursday. I hadn’t missed one yet.
“We need a chatbot” is a proposed solution, not a problem statement. Before any architecture:
Red flags the chapter dramatized: three stakeholders, three different goals, zero numbers; success defined as a technology (“AI”) rather than an outcome; the sponsor who wants the demo before the problem statement.
| They say | You hear | You deliver |
|---|---|---|
| “We need a chatbot” | A solution chosen before the problem | Structured discovery: problem, users, success measures, data, constraints |
| “It must be 100% accurate” | No baseline exists for the current process | Measurable acceptance criteria + error handling proportionate to risk |
| “Innovation” (no metric) | Sponsorship that will dissolve at budget review | Politely ignore until a number attaches |
| “How much will we save?” (too early) | A demand for fiction | The measured baseline now, the projection after feasibility |
| “When it’s wrong, how will I know?” | The end user’s real acceptance criterion | Visible citations + below-threshold flagging — design for when, not if |
Questions drawn from the companion practice-exam bank. Try them before reading the answers section.
During discovery for an AI document-processing system, the operations director insists the system must run fully automated end to end, while the compliance officer states that every output touching customer accounts must be human-reviewed. Both were interviewed separately and each believes their requirement is settled. The project sponsor asks the architect to “just write up the requirements.” What should the architect do FIRST?
An architecture review board challenges an architect’s decision to use a retrieval-based design instead of fine-tuning for a policy-lookup assistant. Two board members favor fine-tuning because a vendor demo impressed them, and the meeting is scheduled to end in a decision. Which TWO actions best serve the architect in communicating the decision?
The remaining 3 questions for this chapter — and every chapter after it — are in the full book, every answer option explained.