Nine hypotheses, three days onsite
Earlier this year I ran a sixty-hour diagnostic of an engineering org I'd never met: three thousand repositories, a hundred-plus Jira projects, five survey instruments, twenty-two interviews, and an AI layer in every phase, because the engagement doubled as an experiment in how much of this work the machines can absorb. The frameworks are public. The useful part was writing nine falsifiable guesses before arriving, and spending the onsite trying to break them.
Before I walked into the building, I had written down nine ways the company might be broken. Named patterns, one line each, and under every one: the evidence that would confirm it, and the evidence that would kill it. Three days onsite is only enough time to do one thing well. I spent them testing.
The nine weren’t hunches — two weeks of system scans and surveys drafted them. All nine survived in some form (the table below keeps the score; one confirmed its way straight into irrelevance), and the belief that died hardest wasn’t on my list at all: the client’s conviction about their own culture.
The engagement was an independent diagnostic my advisory practice ran earlier this year for an established software company, and the client’s question was blunt: why does an org this experienced have this much trouble shipping, and is the technical foundation underneath sound enough to build on once the friction is named? The instruments were DORA, Westrum, and DevEx plus a hypothesis protocol I think matters more than any of them. The engagement also had a second agenda, a deliberate experiment: how much of a traditional org diagnostic can AI carry when you push it into every phase — kickoff, reconnaissance, interviews, synthesis, the readout — instead of bolting it onto one? The two agendas are one design. The machines eat the volume early; the expensive human days go only to pressure-testing the guesses. The fee was structured so I could afford to find out. I have not had that much fun working in years.
#The shape of the engagement
The kickoff brief went out as a document, a ten-minute audio overview, and slides — NotebookLM, one evening, nearly all of it spent writing the source document, which is where the time belongs. 1 The CEO had a long drive ahead that weekend, and the brief rode along; by the time the CEO, senior technology leadership, and I were in a room, everyone had taken in the same plan through whichever medium fit their week.
The shape of it:
| Phase | Effort | What happens |
|---|---|---|
| Pre-work, two weeks remote | ~20 hrs | seven rounds of automated system scans, five survey instruments out, nine hypotheses written, interview grid built |
| Onsite, three days | ~28 hrs | 22 interviews from staff engineer to CFO, architecture walkthroughs, tracing two real customer promises from sale to production, synthesis every evening |
| Synthesis, ten days | ~15 hrs | hypotheses scored, friction register, 30/60/90 plan, board deck, the minisite |
The framing to the client mattered: an operational briefing between people who are about to work together, not an audit and not a sales deck. The CEO commissioned it; engineering leadership knew the scope before I arrived. And since the client is anonymous here: the scores and ratios are exact, the raw inventory counts are rounded, and anything that could identify them is deliberately blurred. That trade is the price of writing about private work at all.
The order of the phases is the argument: everything cheap and asynchronous runs before anything expensive and human. Interview hours were the scarce resource, and I wasn’t going to spend them on facts a script had already pulled.
#The data engine
An org’s own systems will answer questions no interview can, without the performance anxiety. The raw surface: about 140 Jira projects and over three thousand repositories for an engineering group of about sixty — ratios that are findings all by themselves — plus a wiki north of ten thousand pages, a dozen-plus cloud accounts, dozens of incident reviews, and nearly thirty tools in active use.
Platform APIs (Jira, repositories, the wiki, cloud accounts, the user directory) feed collectors running seven scan rounds. Raw JSON dumps are classified and reconciled into twenty-seven CSVs and unified person identities. A generator chain turns the CSVs into 114 chart specifications, which become a 33-page minisite with a source line on every chart. The classified data also feeds data-cited interview questions.
Collection ran in seven rounds before the first conversation, each narrower than the last: a broad inventory pass across every platform, then corrected and expanded metrics, then targeted pulls (reviewer concentration, sprint scope changes, deploy tags), then classification — every Jira project tagged engineering, customer-success, or operations — then content extraction from the wiki’s fifteen key spaces, then a cross-system account audit, and finally a “sponge mode” pass sweeping everything cheap that remained: CI pipeline history, branch-protection audits, thirty sprint retros. About six hours of my attention, because agents did the reading.
The part I’d defend hardest is what happened between collection and use. Raw dumps became twenty-seven classified CSVs; roughly 500 accounts across five systems were reconciled into unified human identities, which is a genuinely hard entity-resolution problem and the foundation for every question about who is load-bearing. From the CSVs, a generator chain wrote 114 chart specifications with one shared house style, and every chart carries a provenance line naming the system it came from and the month it was pulled.
The surveys ran the same weeks: a DORA quick check against the public benchmarks 2 , the six-item Westrum culture instrument 3 , a 23-item DevEx survey across feedback loops, cognitive load, and flow state 4 , the Pragmatic Engineer Test 5 , and an AI-tooling adoption pulse. Nearly sixty people got them; response ran about 70 percent. None of the instruments is proprietary, on purpose: the client should be able to re-run every measurement after I leave and compare.
#Nine named guesses
Six automated system-scan passes and five survey instruments (57 people) feed nine named hypotheses, each written with confirming and killing evidence. Day one narrows the nine to the five with the strongest signal; the other four keep their pre-work evidence and lose their interview slots. Days two and three run targeted interviews against the five, producing scored evidence chains that become the friction register, the 30/60/90 plan, and the leave-alone list.
Every diagnostic framework will happily generate findings forever; writing down what you expect to find, and what would count against it, is what turns a diagnostic into an investigation instead of a tour. Nine patterns, with names ugly enough to argue with:
| Pattern | The guess | What the onsite said |
|---|---|---|
| Portfolio Sprawl | more committed work than demonstrated throughput | confirmed; the headline |
| Sales-Led Chaos | the roadmap re-routes around individual deals | confirmed |
| Fake Platforming | a “platform” that is one product’s internals with an API bolted on | confirmed, and operationally moot |
| Unclear Ownership | formal and operational authority live in different people | confirmed; a reorg had already begun fixing it |
| Dependency Drag | work stalls in queues between teams | confirmed |
| Legacy Gravity | the oldest system taxes every new feature | confirmed |
| Missing Cadence | no reliable operating rhythm between planning moments | confirmed |
| Biz-Tech Mistranslation | the business asks in outcomes, engineering answers in systems | confirmed; the most pervasive |
| Architectural Polarity | two incompatible architectural eras, neither allowed to win | confirmed, weakest signal of the nine |
A scoreboard that clean should make you suspicious. It made me suspicious. The reason it isn’t a horoscope is where the nine came from: the scans and surveys drafted them, so confirmation was the expected outcome by design, and the onsite’s value was pressure — testing each one in rooms with the people best positioned to break it. And the pressure still delivered: it demoted one hypothesis to moot, moved the interview hours to where the evidence was thinnest, and killed the unwritten tenth hypothesis, the one leadership held about the culture. The closest call among the written nine was Architectural Polarity: confirmed, but barely. Half these patterns exist at any company this shape; scoring exists to find the ones that are load-bearing.
Each hypothesis carried a scorecard: confirming evidence, killing evidence, and which interviews could supply either. By the end of day one, nine narrowed to the five with the strongest signal; the other four kept their pre-work evidence and lost their interview slots. Every morning started with the sponsor and the same two questions: what are we trying to prove today, and what are we explicitly not chasing?
#What the data said
| Signal | Reading |
|---|---|
| Epic completion, trailing 12 months | 27.7% of roughly 480 epics |
| Open engineering issues marked “High” priority | 89.4% |
| Work-item lead time, open to done | P50 59 days, P75 210 days |
| PR turnaround in active repos | 27-31 hours average |
| Repositories untouched in 90 days | 75% |
| Repositories with branch protection | 0 — fleet-wide |
| Services with health checks | 10% of about 70 |
| Executive incident reviews with empty preventive-action tables | 3 of 5 |
| Westrum culture mean | 3.82 / 5 — generative |
| DevEx, context-switching item | 2.27 / 5, worst on the survey |
Two pairs of rows carry the engagement. The first pair, 27.7 percent completion against 89.4 percent marked High, says the portfolio is committed to roughly three times what the org demonstrably finishes. The rate is a blunt instrument — epics opened late in the window can’t have finished yet — but blunt doesn’t produce a 3x gap. The lead-time pair says where the time goes: pull requests in active repos turned around in about a day, while the median work item took two months open-to-done, so the delay lives in the queues around engineering, not inside it. That row is Dependency Drag showing up in numbers.
When 89.4 percent of a backlog is marked High, priority is no longer information. It’s wallpaper.
The second surprise is the culture row. Leadership expected Westrum to land near 2.5 — bureaucratic, sliding pathological. It came back 3.82, inside the generative band, with 63 percent of respondents scoring generative. People reported that messengers aren’t shot and failures get treated as learning; what they couldn’t report was finishing anything, because everything was priority one. The most useful sentence in the readout was “your culture is not your problem.” It moved the conversation from culture programs, which take years and rarely work, to portfolio arithmetic, which takes a quarter and does.
The DevEx detail that earns its keep: context switching scored 2.27, the worst item on the survey, which makes it the 90-day canary. If cutting the portfolio works, that number recovers first. If it doesn’t recover, the cut didn’t happen, whatever the roadmap slide says. The re-run is the client’s to fire; the survey package is theirs.
That answers both halves of what the client paid for. Why does shipping hurt: a portfolio committed to three times demonstrated throughput, with the delay living in the queues between teams, not inside engineering. Is the foundation sound: yes, with named exceptions. Review is universal, PRs move in a day, the culture is generative, and one corner of the org already ships the modern way, so the capability exists in house. The exceptions are real — a legacy system that taxes every feature, zero branch protection, health checks on one service in ten — but they are hygiene and age, not rot. That verdict is why the 30/60/90 fixes the portfolio and leaves the architecture alone.
#The interview loop
Twenty-two interviews in three days only works if nothing leaks between sessions, so the capture stack was local by design: a hold-to-talk transcription tool running on-device speech-to-text, with a small local model cleaning the output before it ever hit a file. 6 I told every interviewee exactly that — nothing uploaded, names scrubbed from the transcript afterward — and the promise that goes with it did more for candor than any assurance about my intentions: I don’t play telephone with personal comments; only patterns and blockers get reported.
The cleanup layer is deliberately boring: a small local model with a glossary of about sixty names and product terms, so the transcriber’s phonetic guesses resolve to the right people. Speech models mangle names; a glossary is cheap; wrong names in an evidence chain are not.
Interview prep is where the data engine paid out. Every prepared question cited a specific data point and the hypothesis it tested. Not “tell me about your infrastructure” but: four-fifths of your compute is on previous-generation instances and a fifth of your serverless functions are on end-of-life runtimes — is that a known tradeoff or a surprise? Not “how’s delivery going” but: this board’s completed-issue velocity halved across one quarter — what happened? Questions like that are hard to deflect, and more usefully, they’re respectful: they tell the person you did the reading before spending their hour.
Each evening, the day’s audio re-transcribed overnight on a slower, more accurate model, and the synthesis fan ran: per-person summaries, a hypothesis scorecard update, a contradictions file, an emerging-fixes file, and a what-the-CEO-needs-to-know note. The next morning’s interview guides went out updated with whatever the previous day had surfaced. That rhythm — test all day, consolidate all night, re-aim every morning — is the mechanical heart of the whole method.
A day of interviews is captured locally with names scrubbed. Overnight, a slower and more accurate model re-transcribes the audio, and the synthesis fan runs: per-person notes, a hypothesis scorecard update, a contradictions file. The interview guides are re-aimed at the thinnest evidence, and the loop returns to the next morning's interviews.
For the record, the AI stack also failed onsite: the capture tool silently ate six of Tuesday’s recordings. I ran a phone backup on day one because I assumed something would break; by Tuesday I had stopped assuming, so those six sessions survive only as rough notes. The patterns made it into the synthesis. The verbatim quotes are gone.
#What the client keeps
The readout took two forms: a short operating brief for the executive team and a board deck of eight slides. Behind them, the working artifacts: a friction register of 30 entries, each tiered high, medium, or low, with an owner and an escalation path; the scored hypothesis chains; a decision-rights map for the nine cross-functional decisions that kept surfacing in interviews 7 ; a 30/60/90 plan; and the survey package itself, forms plus scoring script, so every baseline in this post can be re-measured without me.
And then the artifact I’d never shipped before: a private, password-gated minisite — 33 pages, 114 interactive charts, all of it generated from those twenty-seven CSVs, every chart carrying its source-and-date line. The first working version existed three days into pre-work, built from the earliest rounds’ data and regenerated as later rounds landed. It’s a due-diligence-grade appendix the executive team can wander at whatever depth they want, and it changes the readout’s half-life: a PDF ages; a site regenerates from the next data pull.
The section of the plan that got argued about most was none of those. It was the Leave Alone list: the things that looked broken, held up under measurement, and shouldn’t be touched — starting with the culture. An advisor hunting follow-on work does not hand the client a page of things to leave exactly as they are. That page is what keeps a diagnostic from becoming a re-org generator.
#The economics
The premise I modeled the engagement on: the inventory-and-assessment layer of this work — the part that used to justify a large firm sending a team for six months — is now mostly machine work. Sixty hours and a low-five-figure fee bought what the traditional version prices in the six figures and measures in quarters, not because I work faster than a team of ten, but because the volume work stopped requiring humans at all. That’s one mid-size org, with the seller grading his own output, so what I’d defend is the mechanism. AI scanned the repos, indexed the wiki, read every incident review on file, reconciled the identities, drafted the first pass of the register; I did the rooms, the judgment calls, and the narrative.
The client’s own history sharpened the point. This was an executive team that had rejected a several-hundred-page AI-generated report that crossed their desks — and they were right to. AI’s failure mode in consulting is volume cosplaying as rigor. So the delivery constraint here was the opposite: the operating brief reads in ten minutes, every claim traces to a chart or a transcript, and the standing instruction was “if anything feels AI-padded, push back hard.” The CEO also declined daily written recaps — “you just told me; I want to see it at the end” — which is its own kind of discipline. The AI’s job was to make the readout smaller and harder, not longer and softer.
And the data work wasn’t waste, because the working files didn’t die on the consultant’s laptop the way an engagement’s usually do: the collectors, the CSVs, the identity graph, the chart generators, and the survey package all transferred — the engine, not a snapshot.
#The part AI couldn’t do
Reading the room is the obvious thing that didn’t automate. Twenty-two interviews produce maybe six moments that matter — a pause before an answer about ownership, a VP who relaxes when the recorder’s off, two people describing the same meeting as different meetings. The transcript records none of that at the level where it counts. Neither does the synthesis. The night after day one, the most important note I wrote was a private read on what the sponsor actually needed from the engagement — and its design consequence was that the recommendations had to create transparency without adding reporting, because more reporting was the fix the org had already tried. No scan produces that sentence.
The other half is scar tissue. Knowing that 89.4 percent marked-High means triage is dead, that story points populated on 9 percent of stories means specs are theater, that one team’s clean CI pipeline makes it the internal existence-proof the rest of the org should copy — that’s not analysis, it’s recognition, and it comes from years of operating these systems — the director seat I still hold, the companies I built and the ones that broke. The AI found the numbers. Experience is what made specific numbers mean something, and founder-to-CEO trust is what let the meaning land without a defensive war.
Accelerate, Westrum, the DevEx research, the Pragmatic Engineer Test: every instrument here is public, documented, and free, and the AI stack was commodity end to end — transcription, entity resolution, chart generation. That’s the point. Clients don’t pay for frameworks, and they shouldn’t pay for findings-shaped output either. They pay for a stranger disciplined enough to write down how he might be wrong before charging them for being right. The best outcome, and the one this engagement was built around, is that the next diagnostic runs without me.
Notes
- NotebookLM, Google's source-grounded research assistant. Feed it documents; it generates audio overviews, slides, and video summaries that stay inside the sources. The craft is entirely in the source document -- garbage brief, confident garbage podcast. ↩
- The DORA research program, summarized in Forsgren, Humble, and Kim, Accelerate (IT Revolution, 2018), with yearly benchmark updates in the State of DevOps reports. The quick check maps five delivery questions onto the published performance bands. ↩
- Ron Westrum, "A typology of organisational cultures", BMJ Quality & Safety, 2004. Six Likert items type a culture as pathological, bureaucratic, or generative; the instrument travels remarkably well from healthcare to software. Scored here on a 5-point scale with the generative band at 3.5-5.0; DORA fields a 7-point version, so compare bands, not raw means. ↩
- Noda, Storey, Forsgren, and Greiler, "DevEx: What Actually Drives Productivity", ACM Queue, 2023. The three-dimension model (feedback loops, cognitive load, flow state) behind the 23-item survey. ↩
- Gergely Orosz, The Pragmatic Engineer Test: twelve yes/no checks on engineering fundamentals, in the lineage of the Joel Test. The org averaged 6.8 of 12, with code review the one universal bright spot. ↩
- On-device speech-to-text via WhisperKit during sessions; overnight re-transcription with faster-whisper on a larger model with voice-activity filtering. Slower and more accurate is exactly what an overnight window is for. ↩
- The decision-rights mapping uses Bain's RAPID framework. Most of the nine contested decisions had five owners or none; RAPID's job is making that visible enough to be embarrassing. ↩