more deals chased by the same team, with nobody hired.
The real design work was deciding what the AI should leave alone.
Anchor is a platform for producing proposals. Almost every decision that mattered was about what the model was not allowed to do.
Two words first: an RFP is what a client publishes when it wants work done. A pursuit is one attempt at winning one. The firm ran roughly 180 a year.
- Role
- Product & design lead
- Team
- AI · Eng · Architecture · Legal
- Timeline
- Jan 2024–present
- Outcome
- +38% pursuit capacity
Sanitized: no client named, figures rounded, frames rebuilt. Reasoning unchanged.
Every proposal arrives with a rulebook attached.
Page limits per section, mandatory numbering, font floors, attachments due before the narrative.
Reviewers check all of it before reading a word of what you said.
So the last seventy-two hours are the firm’s most expensive people fixing table borders.
And then, sometimes, this happens.
Nobody asked for this. It had to be made countable first.
Proposal pain was accepted as the cost of selling expertise. You cannot get funding for a feeling.
So before any design, I read two years of our proposals against their outcomes as one set. A week of reading turned a complaint into four numbers somebody had to answer for.
A feeling cannot be funded. Four numbers can.
Where is the model allowed to act?
Not “can it write a proposal.” It can. The question is which parts it should be anywhere near.
Every task got sorted by one thing: what happens if it is wrong, and who finds out.
Mechanical, checkable, reversible work goes to the model. Anything a client will hold us to does not.
Consequence rises, authority falls. That single rule is why legal cleared the product in one review, and why the solution architects used it instead of quietly working around it.
The system proposes. A named person owns.
We built a chat interface. Everyone loved it. We deleted it.
The most popular artifact in the program. It took six weeks to work out why it was wrong.
A chat box asks you to already know what to ask. The work is a document with rules attached, so the intelligence belonged in the frame you type into.
- Page budget, live11.3 of 12 pages, 0.7 left. You see the section fill while you can still act on it.
- The rulebook, beside the textNumbering, font floor, reference cap, each holding a state against the section you are in.
- Every reused answer shows its provenanceWho wrote it and when it was last checked, so trust is a decision, not a guess.
Every model call had a price, and the price decided who got to use it.
Running the largest model on everything is the easy build and the expensive habit. What one bid costs to run is a design decision with an invoice attached.
Small cheap model for bulk reading, large one only where judgement is genuinely needed. Same output, a third of the bill.
Routed by whether the task genuinely needs judgment.
At $310 it only paid for itself on the biggest bids. At $92 it ran on all of them.
Nobody puts this in a portfolio. It decided whether the product was real. If only the largest bids can afford to run it, it stays a pilot.
Consulting firms bill clients by the hour. Writing a proposal is not billable — the firm absorbs all 60 senior hours as a bet on winning. So the point was never the saving. It was how many more deals the same people could chase.
Same team. More shots on goal.
bids lost to a formatting or rule failure after launch.
senior hours returned each year to work a client actually pays for.
The headline is not the saving. Cutting hours per bid raised how many bids the same team could run at all. The savings were real; the deals they made room for were worth more.
Four screens, in the order the work happens.
Interactive frames, not flat images — open any of them full size.
Anchor’s production interface is under NDA. The screens on this page were rebuilt for this case study — representative of what shipped, not screenshots of it. The decisions they show, and the numbers beside them, are the real ones.
That’s the story. There’s a great deal more underneath it.
One Firm, One Answer
An AI-first product nobody asked for: what I was solving, how I led it, and where the model was allowed to act.
- Role
- Product & design lead
- Team
- AI · Eng · Architecture · Legal
- Timeline
- Jan 2024–present
- Outcome
- +38% pursuit capacity
A note on this case studyAnchor is internal platform work, published here in sanitized form. No client is named, business figures are rounded or given as ratios, and the interface frames are faithful rebuilds of shipped screens rather than production captures. The research, the reasoning and the design decisions are unchanged.
Three things to know, and one word that carries the rest.
The firm doesn’t sell a product off a shelf. It competes for projects, one at a time. A client publishes an RFP — what it needs, plus the rules for responding — and firms submit written proposals against it.
A pursuit is one attempt at one deal, from the RFP arriving to the proposal going out. The firm ran roughly 180 a year, each costing roughly 60 hours of its most senior people — and not one of those hours is billable.
A big RFP carries thirty to ninety mandatory rules, and the client checks every one before reading your answer. Miss one and you are out, however good the work is. Seven bids went that way in two years.
The client’s document becomes a tracked list of rules on day one.
You write into a frame that will not let you break those rules.
Every line of the estimate shows the past projects behind it.
Wins and losses go back to whoever wrote the answer, by name.
Nobody asked for this product. The first work was not design. It was making the cost countable.
What I owned, what I shaped, and who with.
Make the problem countable before designing anything.
Losses were filed under “price” because nobody writes down that they were thrown out on a page limit. Reading two years of proposals as a single set surfaced seven formatting losses and a 3.1× quality spread between our best and worst version of the same answer. That is what funded the programme.
Draw the autonomy line before drawing a screen.
Every task sorted by what happens if it is wrong. Pulling facts out of a document and finding what we wrote before: yes. A first draft a named human rewrites: yes, clearly labelled. Committing an answer to a client or setting a price: never, however confident the model is. Consequence rises, model authority falls.
Put the intelligence in the frame you type into.
We built the chat interface. It was the most popular artifact in the programme and it was wrong. A chat box asks you to already know what to ask. The work is a document with rules attached, so the page budget runs out in front of you, as you type. Six weeks to learn that, and worth every week.
Own what it costs to run, because that decides who gets to use it.
Which size of model handles which task is a design decision with an invoice attached. Small cheap model for bulk reading, large one only where judgement is genuinely required: $310 down to $92 per bid. At $310 it only paid for itself on 36 bids a year. At $92 it ran on all 180.
Why we lost, through to how the system behaves.
Eight real pieces of the work, labelled by the kind of thinking each one needed. Click any of them to open it full size.
Read, build, price, learn — then round again.
The RFP stops being a PDF.
The 30–90 mandatory rules pulled out on day one rather than discovered in week two: page limits, numbering, minimum font sizes, attachment formats.
30–90 RULES PER DOCUMENTA frame that will not go out of compliance.
The page budget runs down in front of you as you type. Reused answers arrive with an owner and a last-checked date, so reuse stops being a gamble.
3.1× → 1.3× ANSWER SPREADA number that survives procurement.
Every line of the estimate carries the past projects standing behind it, so an architect can defend the number in front of the client’s buyers.
±34% → ±11% OFF THE REAL COSTThe loop finally closes.
Wins and losses return to the answers that earned them, by name. A weak answer that had gone out forty times stopped at forty-one.
RESULTS GO BACK TO THE AUTHORWhere these numbers come from. Baselines came from reading two years of the firm’s own proposals against their recorded outcomes, and from joining estimates to actuals across 60 delivered engagements by hand. Post-launch figures are the firm’s internal pursuit tracking. Absolute values are rounded; the ratios are as measured.
The number worth arguing about is the first one. Proposal work is unbillable — every hour of it is overhead, bet against a chance of winning. So cutting the hours per bid did not mainly save money; it raised how many bids the same team could run at all. The savings were real. The extra deals that capacity opened were worth more.
If you’re hiring at staff or principal level for AI-era product work.
No brief, no request, no budget. The research is what created the problem definition and the funding, before there was a product to design.
Writing down exactly what the AI may do alone, may only draft, and may never touch is why legal cleared it in one review — and why sceptical experts used it instead of working around it.
Choosing which size of model handles which task took the cost of one bid from $310 to $92, which is what turned a tool for the biggest deals into a platform for all 180.
The chat interface was the most popular thing in the programme. Deleting it was the correct call and the hardest one to make politically.
People did not distrust the answer library, they distrusted its age. Showing who wrote each answer, when it was last checked and where it had been sent solved a trust problem with provenance rather than better writing.
Want the reasoning, the research and the screens?
One Firm, One Answer
A services firm wins work by writing proposals. Hundreds a year, every one on a fixed external deadline, every one assembled by the most expensive people in the company. Anchor turned that into a system.
A note on this case studyAnchor is internal platform work, published here in sanitized form. No client is named, business figures are rounded or given as ratios, and the interface frames are faithful rebuilds of shipped screens rather than production captures. The research, the reasoning and the design decisions are unchanged.
Proposals by heroics.
Proposals by system.
A large RFP carries a list of mandatory rules, and the client checks it before reading a word. The firm lost deals it was best qualified to win because a section ran two pages long. The loss reports said price.
I led the research, the shape of the product, and the decisions about where the AI was allowed to act. The rules moved into the page you type in, every answer got a source, and every result came back to whoever wrote it.
Zero bids thrown out on formatting. Forty hours of slack before the deadline instead of six. Senior people back on the work only they can do.
What this is, and why it existed at all.
The firm doesn’t sell a product off a shelf. It competes for projects, one at a time, in writing.
A client publishes an RFP — what it needs, plus the rules for responding. Firms answer with a written proposal. One attempt at one deal is called a pursuit, and the firm ran roughly 180 a year.
Roughly 60 hours of the firm’s most senior people per pursuit, and not one of those hours is billable. It is overhead, bet against a chance of winning.
A large RFP carries thirty to ninety mandatory rules, and the client checks every one before reading your answer. Miss one and you are out, however good the work is.
The client’s document becomes a tracked list of rules on day one.
You write into a frame that will not let you break those rules.
Every line of the estimate shows the past projects behind it.
Wins and losses go back to whoever wrote the answer, by name.
This case study runs on consulting vocabulary. Any term with a dotted underline will explain itself if you hover or tap it, and the glossary below holds all of them in one place.
Nobody read a word of it.
The team moved from a six-hour scramble to forty hours of protected margin, without adding headcount.
Six things were going wrong, and they compounded.
“Slow” was the symptom everyone named. Underneath it sat six failures that fed each other, and the firm had never counted any of them.
The client checks the rules before reading anything. A page limit missed at hour 71 is a zero, and it never appears in a loss report.
The strongest answers lived in people’s heads rather than in a system. When the person who knew one was on a client site, the proposal got a thinner version written that evening.
Answers accurate when written, now citing a metric from an engagement that ended in 2019 or a certification that had lapsed.
Two quiet weeks, then a weekend where the firm’s most expensive people fixed table borders and chased missing biographies. Senior, unpaid, and largely clerical.
Prices came from memory, because the record of what past projects had actually cost was unreachable.
Debriefs were read once and filed. The same weak answer went out roughly forty times over two years.
I read two years of our own proposals.
I sat inside eleven live bids from the RFP arriving to the proposal going out, including three weekends. I interviewed 24 people across solution architecture, bid management, practice leadership, legal and finance — plus four people on the client side who had actually scored our proposals and were willing to say why.
Then I did the thing nobody had done. I read the last two years of submitted proposals as a single set, against whether they won — and joined the pricing system to the delivery system for 60 finished projects, so what we quoted and what it cost could finally sit side by side.
Three rules, and the weekend stopped mattering.
The people scoring our proposals told me their order: rules first, does-it-hang-together second, what-it-actually-says third. Ours was the exact reverse. We wrote the content, argued about the structure, and checked the rules last — at the hour when everyone was tired. Three rules inverted that.
The client’s rulebook becomes the frame.
Page limits, numbering schemes, font floors, and attachment formats are extracted at intake and enforced while people write. Compliance became a property of the document everyone was typing into.
Every answer carries its owner and its date.
A block from the library shows who owns it, when it was last confirmed accurate, where it has been submitted, and how those proposals scored. The same rule governs every line of the estimate.
The outcome comes back to whoever wrote it.
Won, lost, and later the delivery actuals against the original estimate. The library improves because the firm competed, with nobody having to remember to update it.
Read it, build it, price it, then learn from it.
Four parts, one record per bid. The RFP is read once and becomes a set of enforceable rules. The response is written inside those rules. The price cites evidence. The result returns to the library that produced the answer.
Intake
The RFP stops being a PDF and becomes enforceable constraints.
Assemble
A frame that will not let the document go out of compliance.
Price
A number with three prior engagements standing behind it.
Reckon
Outcomes return to the answers that earned them.
Known on day one.
Anchor reads the RFP and produces four things: the rulebook, turned into checks the software can enforce; the list of what the client is asking for; the gaps where we have no ready answer; and the questions worth asking the client before the window for asking closes.
The document argues back.
People write into a frame governed by the client’s own rules. The pages left in a section count down as you type, the way a form shows characters remaining. Text that would push a section past its limit simply will not go in.
This is also where the calendar problem got solved. Open a requirement and Anchor shows the strongest answer the firm has ever submitted to that question, ranked by how those proposals scored, with who wrote it and when it was last checked attached. The author does not need to be free that week. Their best work is.
A number that survives procurement.
Four things make up a price: the tasks the work breaks into, who staffs them, the margin, and money set aside for what might go wrong. Every line cites three comparable past projects and what they actually cost. The estimate comes back as a range, along with the reasons the range is wide.
The research found the gap between quote and reality clustered in four kinds of work — and in all four, architects had raised the risk out loud in review and then rounded it away. The template had nowhere to put uncertainty, so uncertainty turned into false precision at the last step before submission. Giving it somewhere to live is most of the fix.
Forty uses of a losing answer stopped at forty-one.
After submission: won, lost, withdrawn, and the client’s own scoring where they share it. Months later the real project costs arrive and grade the original estimate. Both loops report back to the people who produced the work, by name.
More consequenceless model authority
Where the model was allowed to act.
The most consequential design work on Anchor was deciding, task by task, where the AI could act on its own, where it could only propose something for a person to approve, and where it had no path at all. That one decision made the product safer, cheaper and more trusted at the same time.
small model
ratified by a person
to a client
Work with a right answer a human can check in seconds. The model runs it unattended and logs what it did.
Work where the model is useful and fallible. It produces something to argue with, and a named person ratifies it before it moves.
Commitments that bind the firm to a client. Kept out by design rather than by capability, so there is nothing to disable.
Two models, split by whether the task needs judgment.
Treating every call as equally hard is how these programs quietly become unaffordable. I split the workload by task class and priced each one, then held the split as a design constraint rather than an optimization to do later.
Extraction, classification, retrieval, formatting checks, freshness scoring. Work with a verifiable right answer, where the expensive model measured no better in evaluation.
Solution shape, work breakdown, pricing rationale, clarification questions. Work where the output is an argument rather than a fact.
What it costs to run one bid through the AI. That figure changed who could use it, not just the budget: cheap enough for all 180 bids a year rather than the largest fifth — and the small bids were exactly where the formatting losses had been happening.
A hallucinated page limit blocks writing that should have been allowed. Loud, annoying, and safe.
Every rule it finds links back to the sentence and page it came from. One click dismisses it with a reason, and that dismissal becomes training data. A false alarm costs seconds, so we deliberately tuned it to over-report.
The dangerous one. Silent, invisible until the client rejects the submission, and the exact failure Anchor existed to prevent.
Mandatory sections get a second pass from a different model. Where the two disagree, a person reads the source page. Double-checking only where the cost of being wrong is catastrophic; a single pass everywhere else.
A confident number resting on the wrong evidence. This one damages trust fastest, because it is only discovered in a client room.
Citations show why they were matched, in plain language. Architects reject a citation in one click and the estimate recomputes without it. The band widens honestly instead of the number quietly holding.
Five calls, and what each one cost or returned.
The screens were the easy part. These are the decisions that determined whether any of it worked — including the two that made people unhappy.
Compliance can block submission.
A checklist owned by the bid manager.
Checklists are weakest in the final 24 hours, the moment every recorded miss occurred.
Seven pursuits were disqualified even though every team had a checklist.
format disqualifications. Each avoided one is worth the full sunk cost of a pursuit plus the expected value of the contract.
No chat interface.
A sidebar copilot that demoed well.
Conversation adds turns when teams are already in hour sixty.
Chat usage collapsed after two weeks. In-page drafts held.
to a usable first draft. More importantly, usage held after the novelty period, which the chat prototype had not.
Small model for volume. Large model only for judgment.
One frontier model for every call.
For 85% of the workload, the expensive model measured no better.
The lower routing cost made coverage of all 180 pursuits viable.
per pursuit. That changed coverage rather than only cost: Anchor became cheap enough to run on all 180 pursuits instead of the largest fifth, and the small ones were where the disqualifications had been.
The system proposes. A named person owns.
Full auto-assembly with review afterward.
A submitted answer is a commitment the firm has made.
Every ratification records who stood behind the claim.
Legal cleared the product in a single pass instead of becoming a launch blocker. Senior architects, the group most likely to reject the tool, adopted it ahead of everyone else.
Win and loss data goes back to contributors, by name.
Keep outcomes with pursuit leadership.
A library improves only when authors see what their work did.
One losing answer had circulated for two years without its authors knowing.
The largest single lift in answer quality landed in the two quarters after this went live, before any model change shipped.
Seventy-two hours, and the last six decided it.
Two quiet weeks, then a weekend. Principals reformatting sections at 2am, compliance checked last, the price set by whoever was most confident in the room.
Compliance at the frame. Evidence behind the number.
Constraints enforced while people wrote. The strongest existing answer surfaced regardless of who was staffed. Every price line standing on three prior engagements.
More shots on goal, same team.
Proposal work is unbillable, so every hour it consumes is overhead charged against a chance of winning. Cutting the hours did more than save money. It raised how many pursuits the firm could run at all.
Submitted with time to spare, and compliant by construction.
Three surfaces.
Three moments the system became real.
Specific moments when people could see time, evidence, or their own best work differently.
“We passed on a pursuit on day one that we would otherwise have chased for three weeks. The brief said we had no answer for two mandatory requirements and no path to one before the deadline. That was the first time declining felt like a decision rather than a defeat.”
Practice Lead, Data & Analytics
“I watched the page budget run out on a section and the block simply would not take more text. Eleven years of proposals and that is the first time the document argued back. We submitted a day early.”
Bid Manager
“Procurement pushed on our estimate for forty minutes. I opened the panel and walked them through three prior engagements and what each one actually cost. They stopped pushing. That has never happened to me before.”
Principal Solution Architect
pursuit capacity, at the same headcount
Planning estimate: ~6,800 hours × $150 fully loaded senior-hour. This is capacity value; none of it was booked as cash savings.
There are shorter versions of this same project.
One surface, and the decisions inside it.
The workspace an architect actually writes in, rebuilt here at full fidelity. Everything on this page argues about where the model should stop; this is what that argument looks like as an interface.
- The core decisionThe rules live in the frame, not in a chat box.
The chat interface tested well and was deleted. An architect at hour sixty does not want to ask a question; they want something already on the page to argue with. The page budget and the rulebook hold state beside the text you are writing.
- Budget11.3 of 12 pages, counting down live.
The constraint that disqualifies you is shown while you can still act on it. Told at submission, it is a post-mortem; told as you type, it is a design.
- TrustEvery suggested answer carries its provenance.
Who wrote it, when it was last checked, where it has been sent. People did not distrust the answer library — they distrusted its age. So age became a visible property rather than a hidden one.
- AuthorshipThe model proposes; a named person decides.
Nothing the model produces is styled as finished. Drafted content is visually provisional until a human commits it, because the interface has to make the autonomy line legible without anyone explaining it.