Every tool advances through a defined authorization path with named approvers — the relevant operating division, Counsel, Privacy, and Disclosure — and aligns to IRM, IRC, and CFR authority before any field release. A human examiner reviews every output; the tool never makes a determination, only surfaces an indication to review. We treat the speed and clarity of approvals as the source of leadership and field trust, not an obstacle to it.
The right people sign off before any tool is used, and it lines up with the rules. A person reviews everything the tool produces — the tool only points things out, it never decides. Fast, clear approvals are what earn trust, not red tape.
Tools operate read-only within an already-authorized boundary, reading from ECM, IDRS, AIMS, and RGS and writing back to none. We plan against Authority to Operate, Pub 1075 FTI handling, and PCLIA requirements as gating checkpoints, not afterthoughts. PII and FTI controls, and a clear answer to where data lives during processing, are settled before any tool reaches a workstation.
Tools only read the official case systems — they never change them. We clear the security and privacy sign-offs up front, and we settle exactly where information lives before anything reaches a desk. No real taxpayer data is ever used.
The best resource is the front-line workforce. Tools are designed and tested directly with a small group of working examiners and revenue officers throughout development — they are the source of tools that actually get used. That direct, lightweight engagement during the build matters more than broad committee input; wider coordination is sequenced to the pre-launch gate.
The people who do the work help build the tools, hands-on, the whole way through — they're the reason the tools end up useful. The wider review across offices happens before launch, not during the build, so the work stays fast.
The three tracks — SB/SE, LB&I, and Collections — build on one shared foundation rather than three disconnected tools. That common base keeps outputs consistent and de-duplicates each tool against the existing system of record, so the workbench complements ECM, IDRS, AIMS, and RGS rather than competing with them. Cross-division coordination is sequenced to the pre-launch gate.
All three areas build on one shared base instead of three separate tools, so everything stays consistent and nothing duplicates the official systems. The broad coordination across offices happens before launch — at the right time, not during the build.
The plan is deliberately bounded: a defined slate sequenced foundation-first, with the genuinely hard, judgment-heavy issues placed late rather than over-promised early. Anything outside the window is ticketed, not crammed in. We would rather prove one fully validated tool — held to a defined accuracy standard and Counsel review — and earn the rest.
We do a few high-value things well rather than everything at once. The genuinely hard items come later; the rest waits its turn. Better to prove one solid tool and earn the next than over-promise a long list.
Within ninety days we can credibly measure delivery, adoption, and output integrity, including citation-verification rates and examiner-rated usefulness. Downstream effects on cycle time and case quality are the value hypothesis these leading indicators are designed to test, not outcomes promised over a pilot. Baselines and targets are set with the field after the pilot.
We measure what we can honestly show in 90 days — did people use it, did it hold up. The bigger payoff, like faster cases, is what we're testing toward, not something we promise up front.
| Risk | Mitigation |
|---|---|
| The Day-0 workspace and authorizations slip; a two-week delay does not cost two weeks — it compresses the harder work at the end of the window.The approved workspace and the sign-offs to start are delayed. | → Settle the environment plus three named gates in week one — security authorization, synthetic-data/Privacy sign-off, and Counsel/Disclosure redaction review — each with an owner. This is the critical path.Settle them in the first week and put one person in charge of each — this is the most likely thing to break the schedule. |
| Eighteen deliverables is a count, not a commitment; several encode contested domain logic that eats the calendar.Promising all the tools at once over-reaches. | → Re-tier into a committed core (the shared foundation plus a clean subset), with the rest targeted or stretch — under-promise and beat it; ticket additions to a later cycle.Commit to a confident core, and label the harder tools as stretch goals — under-promise, over-deliver. |
| Field adoption lags because examiners are case-cycle bound and have little spare capacity.Busy staff don't have time to pick up new tools. | → Lead design with working examiners, pilot a small cohort, plan dedicated training time, and let usefulness drive voluntary use.Build them with the staff, start with a small group, set aside training time, and let usefulness — not a mandate — drive use. |
| An inaccurate citation or date is inherited by the examiner and reaches the taxpayer.A wrong source or date slips through to the taxpayer. | → Treat accuracy as a release gate: every output is pin-cited, a defined validation standard and Counsel / Practice Area sign-off precede shipping, and corrections are tracked.Don't ship a tool until its accuracy is checked and signed off; every answer shows its source, and fixes are tracked. |
| The redaction core misses something and exposes sensitive detail — the single highest-blast-radius failure.The tool that hides sensitive details misses one. | → Treat redaction as assistive and human-verified — it suggests; a person clears every page; it never marks content safe on its own.It only suggests — a person checks and clears every page by hand. It can never mark a page safe on its own. |
| Tools tuned on synthetic data encode wrong assumptions about messy real records (scanning noise, missing years, malformed forms).Tools built on clean made-up data stumble on messy real records. | → Run an explicit synthetic-to-real validation step before any tool leaves the pilot, and state the gap openly rather than hide it.Test deliberately against real-world mess before any tool leaves practice mode, and say so openly. |
| Over-reliance lets a tool's output read as a determination rather than a reviewed indication.People start treating a tool's answer as the decision. | → Embed indication-not-determination, read-only, post-selection, and human-in-the-loop in every output, with a review log on each result.Every result is clearly marked as something to review, never the decision, with a person in the loop and a record kept. |
| Over-reliance on a single contractor for the environment, AI assist, or plumbing is a single point of failure on the critical path — and no one is set up to maintain the tools after Day 90.Too much rides on one outside vendor, and no one owns upkeep after 90 days. | → Anchor on field-built, agency-owned tooling; the IRS receives all source and AI-assisted code with clear ownership and licensing, and a named maintenance owner is in place before the pilot ends.Keep the work and the code in government hands with clear ownership, and name who maintains it before the pilot ends. |
The first decision of the pilot is not what we build, but where it can legitimately be built. This is a software-delivery effort, and software delivery runs on a standard toolchain — version control, issue tracking, and build/deploy automation, alongside AI-assisted development. A standard compliance-issue laptop is correctly configured for casework, not for building and shipping software — a known provisioning constraint, not a gap in IT. All three paths share the same guardrails: synthetic or sanitized data only, no Federal Tax Information ever in the environment, and the IRS owning every deliverable.The first decision isn't what we build — it's where we're allowed to build it. Building software needs proper software tools (a place to manage the code, track the work, and test it), and a standard casework laptop isn't set up for that — by design, and that's fine. Whatever we use: made-up data only, no real taxpayer information, and the government owns everything we build.
Six tools per track, one common spine. Each tool is a thin, governed layer on a shared base — so output stays consistent, every tool de-duplicates against the system of record, and the marginal cost of the next tool falls.Six tools per area, but one shared base underneath. Each tool is a thin layer on top, so they stay consistent and the next one costs less to build.
In ninety days we can credibly measure delivery, adoption, and output integrity. Downstream effects on cycle time and case quality are the hypothesis these indicators are designed to test — baselines and targets set with the field after the pilot, never pre-promised.What we'll actually measure in the first 90 days — early signals, not promised outcomes. The bigger payoff is what these are testing toward, set with the field after the pilot.