↳CASE STUDY · KELLTON · REVENUE OPERATIONS

One Firm, One Answer

A services firm wins work by writing proposals. Hundreds a year, every one on a fixed external deadline, every one assembled by the most expensive people in the company. Anchor turned that into a system.

ROLEProduct & design lead · AI-first
SCOPE4 surfaces · pursuit platform
TEAMApplied AI · Eng · Solution architecture · Legal
IMPACT+38% pursuit capacity
How much time do you have?

A note on this case studyAnchor is internal platform work, published here in sanitized form. No client is named, business figures are rounded or given as ratios, and the interface frames are faithful rebuilds of shipped screens rather than production captures. The research, the reasoning and the design decisions are unchanged.

THE 60-SECOND READ

Proposals by heroics.
Proposals by system.

01 · ChallengeWe were losing on formatting.

A large RFP carries a list of mandatory rules, and the client checks it before reading a word. The firm lost deals it was best qualified to win because a section ran two pages long. The loss reports said price.

02 · My leadershipA frame, a source, a loop.

I led the research, the shape of the product, and the decisions about where the AI was allowed to act. The rules moved into the page you type in, every answer got a source, and every result came back to whoever wrote it.

03 · Outcome38% more deals. Same team.

Zero bids thrown out on formatting. Forty hours of slack before the deadline instead of six. Senior people back on the work only they can do.

BEFORE WE START · THE SETUP

What this is, and why it existed at all.

The firm doesn’t sell a product off a shelf. It competes for projects, one at a time, in writing.

01 · How the work arrives

A client publishes an RFP — what it needs, plus the rules for responding. Firms answer with a written proposal. One attempt at one deal is called a pursuit, and the firm ran roughly 180 a year.

02 · What each one costs

Roughly 60 hours of the firm’s most senior people per pursuit, and not one of those hours is billable. It is overhead, bet against a chance of winning.

03 · Why it kept failing

A large RFP carries thirty to ninety mandatory rules, and the client checks every one before reading your answer. Miss one and you are out, however good the work is.

Solution architectsThe senior technical people who design what the firm proposes and work out what it should cost. My primary users.
Bid managementThe people who run a submission: the schedule, the rulebook, and getting the document out before the deadline.
Practice leadershipThe people who decide which deals are worth chasing at all, and what team capacity really means.
LegalWho decide what an AI system may and may not touch. They cleared this one in a single review.
Intakereading the rules

The client’s document becomes a tracked list of rules on day one.

Assemblewriting inside them

You write into a frame that will not let you break those rules.

Pricedefending the number

Every line of the estimate shows the past projects behind it.

Reckonlearning from the result

Wins and losses go back to whoever wrote the answer, by name.

A note on the language

This case study runs on consulting vocabulary. Any term with a dotted underline will explain itself if you hover or tap it, and the glossary below holds all of them in one place.

email attachment pasted from chat 2021 proposal SME draft v3 bio.docx security answer, author unknown pricing_final_v9.xlsx client template, section 4
HOURS TO SUBMISSION72
One pursuit. One deadline.
4.1
4.2
4.3
4.4
Disqualified. Section 4.2 exceeded the stated page limit.

Nobody read a word of it.

→REDESIGNED
34hours of deadline margin reclaimed

The team moved from a six-hour scramble to forty hours of protected margin, without adding headcount.

THE PROBLEM

Six things were going wrong, and they compounded.

“Slow” was the symptom everyone named. Underneath it sat six failures that fed each other, and the firm had never counted any of them.

01
Losing on format.

The client checks the rules before reading anything. A page limit missed at hour 71 is a zero, and it never appears in a loss report.

02
The best answer was hostage to a calendar.

The strongest answers lived in people’s heads rather than in a system. When the person who knew one was on a client site, the proposal got a thinner version written that evening.

03
The library had quietly rotted.

Answers accurate when written, now citing a metric from an engagement that ended in 2019 or a certification that had lapsed.

04
The last seventy-two hours.

Two quiet weeks, then a weekend where the firm’s most expensive people fixed table borders and chased missing biographies. Senior, unpaid, and largely clerical.

05
Five architects, five prices.

Prices came from memory, because the record of what past projects had actually cost was unreachable.

06
Nothing was learned from a loss.

Debriefs were read once and filed. The same weak answer went out roughly forty times over two years.

3.1×Spread between our strongest and weakest version of the same answer
7Bids thrown out on formatting in the same window
~60Senior hours per bid, almost none of it paid for
±34%Gap between what we quoted and what it cost
THE RESEARCH

I read two years of our own proposals.

I sat inside eleven live bids from the RFP arriving to the proposal going out, including three weekends. I interviewed 24 people across solution architecture, bid management, practice leadership, legal and finance — plus four people on the client side who had actually scored our proposals and were willing to say why.

Then I did the thing nobody had done. I read the last two years of submitted proposals as a single set, against whether they won — and joined the pricing system to the delivery system for 60 finished projects, so what we quoted and what it cost could finally sit side by side.

WHAT THE RESEARCH CHANGED

Three rules, and the weekend stopped mattering.

The people scoring our proposals told me their order: rules first, does-it-hang-together second, what-it-actually-says third. Ours was the exact reverse. We wrote the content, argued about the structure, and checked the rules last — at the hour when everyone was tired. Three rules inverted that.

01 · CONFORM

The client’s rulebook becomes the frame.

Page limits, numbering schemes, font floors, and attachment formats are extracted at intake and enforced while people write. Compliance became a property of the document everyone was typing into.

02 · SOURCE

Every answer carries its owner and its date.

A block from the library shows who owns it, when it was last confirmed accurate, where it has been submitted, and how those proposals scored. The same rule governs every line of the estimate.

03 · RETURN

The outcome comes back to whoever wrote it.

Won, lost, and later the delivery actuals against the original estimate. The library improves because the firm competed, with nobody having to remember to update it.

ONE CONNECTED SYSTEM

Read it, build it, price it, then learn from it.

Four parts, one record per bid. The RFP is read once and becomes a set of enforceable rules. The response is written inside those rules. The price cites evidence. The result returns to the library that produced the answer.

Read

Intake

The RFP stops being a PDF and becomes enforceable constraints.

The rulebook
Build

Assemble

A frame that will not let the document go out of compliance.

Rules → draft
Price

Price

A number with three prior engagements standing behind it.

Draft → priced
Learn

Reckon

Outcomes return to the answers that earned them.

Priced → learned
Read
INTAKE · READING THE CLIENT’S RULES

Known on day one.

Anchor reads the RFP and produces four things: the rulebook, turned into checks the software can enforce; the list of what the client is asking for; the gaps where we have no ready answer; and the questions worth asking the client before the window for asking closes.

The Q&A window is a real lever. Most firms miss it because nobody has read the document closely enough in the first 48 hours to know what to ask. We started asking on day one.
Bid or no bid, on evidence. The most valuable decision in pursuit management is declining early. It used to happen on instinct in week two.
30–90MANDATORY CONSTRAINTS EXTRACTED PER RFP
Intake summary, the compliance matrix, and the gap view
Build
ASSEMBLE · WRITING INSIDE THE RULES
3.1× → 1.3×SPREAD BETWEEN OUR BEST AND WORST ANSWER

The document argues back.

People write into a frame governed by the client’s own rules. The pages left in a section count down as you type, the way a form shows characters remaining. Text that would push a section past its limit simply will not go in.

This is also where the calendar problem got solved. Open a requirement and Anchor shows the strongest answer the firm has ever submitted to that question, ranked by how those proposals scored, with who wrote it and when it was last checked attached. The author does not need to be free that week. Their best work is.

Provenance on every block. Owner, last verified, where it has been sent, how it scored.
A voice pass. Register normalized across contributors without flattening the technical content. Evaluators score coherence, and a document that reads like five companies loses points.
The workspace, the library, the best-answer ranking, the voice pass, the submission gate, and provenance
Price
PRICE · A NUMBER THAT SHOWS ITS WORKING

A number that survives procurement.

Four things make up a price: the tasks the work breaks into, who staffs them, the margin, and money set aside for what might go wrong. Every line cites three comparable past projects and what they actually cost. The estimate comes back as a range, along with the reasons the range is wide.

The research found the gap between quote and reality clustered in four kinds of work — and in all four, architects had raised the risk out loud in review and then rounded it away. The template had nowhere to put uncertainty, so uncertainty turned into false precision at the last step before submission. Giving it somewhere to live is most of the fix.

Why this number. A panel written to be read aloud in a client room rather than to sit in an appendix.
The band is the point. Architects trusted a number that admitted what it did not know far more readily than one that did not.
±34% → ±11%ESTIMATE VARIANCE AGAINST DELIVERED ACTUALS
Work breakdown with citations, and the why-this-number panel
Learn
RECKON · LEARNING FROM THE RESULT

Forty uses of a losing answer stopped at forty-one.

After submission: won, lost, withdrawn, and the client’s own scoring where they share it. Months later the real project costs arrive and grade the original estimate. Both loops report back to the people who produced the work, by name.

Answers carry a record. Every block accumulates a history of how proposals containing it performed.
Estimates get graded. The gap between quote and reality stopped being an annual finance exercise and became feedback the architect could actually see.
The scorecard, the bid brief, and the requirement list
THE AUTONOMY LINE

Where the model was allowed to act.

The most consequential design work on Anchor was deciding, task by task, where the AI could act on its own, where it could only propose something for a person to approve, and where it had no path at all. That one decision made the product safer, cheaper and more trusted at the same time.

Model actscheckable work
small model
Constraint extractionRequirement parsingFormat validationPage budget mathLibrary retrievalFreshness flagging
On the linedrafted by the model
ratified by a person
Solution shapeWork breakdownPricing rationaleClarification questionsVoice normalizationComparable matching
Human onlya commitment
to a client
Setting the priceRatifying an answerDeclining a bidOverriding complianceNaming staff
Above the line

Work with a right answer a human can check in seconds. The model runs it unattended and logs what it did.

On the line

Work where the model is useful and fallible. It produces something to argue with, and a named person ratifies it before it moves.

Below the line

Commitments that bind the firm to a client. Kept out by design rather than by capability, so there is nothing to disable.

Two models, split by whether the task needs judgment.

Treating every call as equally hard is how these programs quietly become unaffordable. I split the workload by task class and priced each one, then held the split as a design constraint rather than an optimization to do later.

Small model handles

Extraction, classification, retrieval, formatting checks, freshness scoring. Work with a verifiable right answer, where the expensive model measured no better in evaluation.

Large model handles

Solution shape, work breakdown, pricing rationale, clarification questions. Work where the output is an argument rather than a fact.

Before$310
After$92

What it costs to run one bid through the AI. That figure changed who could use it, not just the budget: cheap enough for all 180 bids a year rather than the largest fifth — and the small bids were exactly where the formatting losses had been happening.

WHEN THE MODEL IS WRONG
Failure 01 It extracts a constraint that is not there.

A hallucinated page limit blocks writing that should have been allowed. Loud, annoying, and safe.

What we built

Every rule it finds links back to the sentence and page it came from. One click dismisses it with a reason, and that dismissal becomes training data. A false alarm costs seconds, so we deliberately tuned it to over-report.

Failure 02 It misses a constraint entirely.

The dangerous one. Silent, invisible until the client rejects the submission, and the exact failure Anchor existed to prevent.

What we built

Mandatory sections get a second pass from a different model. Where the two disagree, a person reads the source page. Double-checking only where the cost of being wrong is catastrophic; a single pass everywhere else.

Failure 03 It cites an engagement that is not comparable.

A confident number resting on the wrong evidence. This one damages trust fastest, because it is only discovered in a client room.

What we built

Citations show why they were matched, in plain language. Architects reject a citation in one click and the estimate recomputes without it. The band widens honestly instead of the number quietly holding.

THE DECISIONS THAT PAID

Five calls, and what each one cost or returned.

The screens were the easy part. These are the decisions that determined whether any of it worked — including the two that made people unhappy.

01

Compliance can block submission.

Instead

A checklist owned by the bid manager.

Why

Checklists are weakest in the final 24 hours, the moment every recorded miss occurred.

Proof

Seven pursuits were disqualified even though every team had a checklist.

7 → 0

format disqualifications. Each avoided one is worth the full sunk cost of a pursuit plus the expected value of the contract.

02

No chat interface.

Instead

A sidebar copilot that demoed well.

Why

Conversation adds turns when teams are already in hour sixty.

Proof

Chat usage collapsed after two weeks. In-page drafts held.

40 min → 4 min

to a usable first draft. More importantly, usage held after the novelty period, which the chat prototype had not.

03

Small model for volume. Large model only for judgment.

Instead

One frontier model for every call.

Why

For 85% of the workload, the expensive model measured no better.

Proof

The lower routing cost made coverage of all 180 pursuits viable.

$310 → $92

per pursuit. That changed coverage rather than only cost: Anchor became cheap enough to run on all 180 pursuits instead of the largest fifth, and the small ones were where the disqualifications had been.

04

The system proposes. A named person owns.

Instead

Full auto-assembly with review afterward.

Why

A submitted answer is a commitment the firm has made.

Proof

Every ratification records who stood behind the claim.

One review

Legal cleared the product in a single pass instead of becoming a launch blocker. Senior architects, the group most likely to reject the tool, adopted it ahead of everyone else.

05

Win and loss data goes back to contributors, by name.

Instead

Keep outcomes with pursuit leadership.

Why

A library improves only when authors see what their work did.

Proof

One losing answer had circulated for two years without its authors knowing.

Two quarters

The largest single lift in answer quality landed in the two quarters after this went live, before any model change shipped.

ONE PURSUIT, BEFORE AND AFTER
The old shape of a pursuit

Seventy-two hours, and the last six decided it.

Two quiet weeks, then a weekend. Principals reformatting sections at 2am, compliance checked last, the price set by whoever was most confident in the room.

The system

Compliance at the frame. Evidence behind the number.

Constraints enforced while people wrote. The strongest existing answer surfaced regardless of who was staffed. Every price line standing on three prior engagements.

The operating result

More shots on goal, same team.

Proposal work is unbillable, so every hour it consumes is overhead charged against a chance of winning. Cutting the hours did more than save money. It raised how many pursuits the firm could run at all.

WHAT THE FIRM SAW

Three surfaces.
Three moments the system became real.

Specific moments when people could see time, evidence, or their own best work differently.

01 INTAKE · THE DECISION TO DECLINE

“We passed on a pursuit on day one that we would otherwise have chased for three weeks. The brief said we had no answer for two mandatory requirements and no path to one before the deadline. That was the first time declining felt like a decision rather than a defeat.”

Practice Lead, Data & Analytics
02 ASSEMBLE · THE FRAME HELD

“I watched the page budget run out on a section and the block simply would not take more text. Eleven years of proposals and that is the first time the document argued back. We submitted a day early.”

Bid Manager
03 PRICE · THE NUMBER SURVIVED THE ROOM

“Procurement pushed on our estimate for forty minutes. I opened the panel and walked them through three prior engagements and what each one actually cost. They stopped pushing. That has never happened to me before.”

Principal Solution Architect
Anchor shipped with the same accessibility standard as everything else I lead: keyboard-complete, documented focus order, and no state signalled by color alone.
Annual operating impact +38%

pursuit capacity, at the same headcount

~6,800principal hours returned each year
0pursuits lost to format
≈$1.0Mestimated senior-labor value returned

Planning estimate: ~6,800 hours × $150 fully loaded senior-hour. This is capacity value; none of it was booked as cash savings.

There are shorter versions of this same project.

THE CRAFT, UP CLOSE

One surface, and the decisions inside it.

The workspace an architect actually writes in, rebuilt here at full fidelity. Everything on this page argues about where the model should stop; this is what that argument looks like as an interface.

Anchor · the writing workspace · rebuilt from the shipped screen
  • The core decisionThe rules live in the frame, not in a chat box.

    The chat interface tested well and was deleted. An architect at hour sixty does not want to ask a question; they want something already on the page to argue with. The page budget and the rulebook hold state beside the text you are writing.

  • Budget11.3 of 12 pages, counting down live.

    The constraint that disqualifies you is shown while you can still act on it. Told at submission, it is a post-mortem; told as you type, it is a design.

  • TrustEvery suggested answer carries its provenance.

    Who wrote it, when it was last checked, where it has been sent. People did not distrust the answer library — they distrusted its age. So age became a visible property rather than a hidden one.

  • AuthorshipThe model proposes; a named person decides.

    Nothing the model produces is styled as finished. Drafted content is visually provisional until a human commits it, because the interface has to make the autonomy line legible without anyone explaining it.

NEXT CASE STUDY Amazon. The Talent Ecosystem →