Version
Roadmap Classic v2.3 ✦ Enhanced
Working draft · proposed implementation plan · for leadership discussion · not yet an official IRS commitment or schedule

Compliance Workbench — 90-Day Implementation Roadmap

SB/SE Exam | LB&I Exam | Collections — targeted deliverables · Q3 2026 (proposed)
Overall
0%

Executive considerations

This roadmap sequences a focused, 90-day rollout of practical compliance tools across three functions. It is a proposed plan for discussion — the deliverables below are starting points to refine with leadership, and production timing remains subject to the agency's existing governance, security, and resourcing processes.A 90-day plan to give frontline staff practical tools. It's a starting point for discussion, not a final commitment — and a real rollout still depends on the agency's normal approvals.

Governance & Approvals

Every tool advances through a defined authorization path with named approvers — the relevant operating division, Counsel, Privacy, and Disclosure — and aligns to IRM, IRC, and CFR authority before any field release. A human examiner reviews every output; the tool never makes a determination, only surfaces an indication to review. We treat the speed and clarity of approvals as the source of leadership and field trust, not an obstacle to it.

The right people sign off before any tool is used, and it lines up with the rules. A person reviews everything the tool produces — the tool only points things out, it never decides. Fast, clear approvals are what earn trust, not red tape.

Authorization & Data

Tools operate read-only within an already-authorized boundary, reading from ECM, IDRS, AIMS, and RGS and writing back to none. We plan against Authority to Operate, Pub 1075 FTI handling, and PCLIA requirements as gating checkpoints, not afterthoughts. PII and FTI controls, and a clear answer to where data lives during processing, are settled before any tool reaches a workstation.

Tools only read the official case systems — they never change them. We clear the security and privacy sign-offs up front, and we settle exactly where information lives before anything reaches a desk. No real taxpayer data is ever used.

Front-Line-Led Development

The best resource is the front-line workforce. Tools are designed and tested directly with a small group of working examiners and revenue officers throughout development — they are the source of tools that actually get used. That direct, lightweight engagement during the build matters more than broad committee input; wider coordination is sequenced to the pre-launch gate.

The people who do the work help build the tools, hands-on, the whole way through — they're the reason the tools end up useful. The wider review across offices happens before launch, not during the build, so the work stays fast.

Shared Foundation, One System of Record

The three tracks — SB/SE, LB&I, and Collections — build on one shared foundation rather than three disconnected tools. That common base keeps outputs consistent and de-duplicates each tool against the existing system of record, so the workbench complements ECM, IDRS, AIMS, and RGS rather than competing with them. Cross-division coordination is sequenced to the pre-launch gate.

All three areas build on one shared base instead of three separate tools, so everything stays consistent and nothing duplicates the official systems. The broad coordination across offices happens before launch — at the right time, not during the build.

Focused Scope, Real Value

The plan is deliberately bounded: a defined slate sequenced foundation-first, with the genuinely hard, judgment-heavy issues placed late rather than over-promised early. Anything outside the window is ticketed, not crammed in. We would rather prove one fully validated tool — held to a defined accuracy standard and Counsel review — and earn the rest.

We do a few high-value things well rather than everything at once. The genuinely hard items come later; the rest waits its turn. Better to prove one solid tool and earn the next than over-promise a long list.

Measurement & Value

Within ninety days we can credibly measure delivery, adoption, and output integrity, including citation-verification rates and examiner-rated usefulness. Downstream effects on cycle time and case quality are the value hypothesis these leading indicators are designed to test, not outcomes promised over a pilot. Baselines and targets are set with the field after the pilot.

We measure what we can honestly show in 90 days — did people use it, did it hold up. The bigger payoff, like faster cases, is what we're testing toward, not something we promise up front.

Key risks & mitigations

RiskMitigation
The Day-0 workspace and authorizations slip; a two-week delay does not cost two weeks — it compresses the harder work at the end of the window.The approved workspace and the sign-offs to start are delayed. Settle the environment plus three named gates in week one — security authorization, synthetic-data/Privacy sign-off, and Counsel/Disclosure redaction review — each with an owner. This is the critical path.Settle them in the first week and put one person in charge of each — this is the most likely thing to break the schedule.
Eighteen deliverables is a count, not a commitment; several encode contested domain logic that eats the calendar.Promising all the tools at once over-reaches. Re-tier into a committed core (the shared foundation plus a clean subset), with the rest targeted or stretch — under-promise and beat it; ticket additions to a later cycle.Commit to a confident core, and label the harder tools as stretch goals — under-promise, over-deliver.
Field adoption lags because examiners are case-cycle bound and have little spare capacity.Busy staff don't have time to pick up new tools. Lead design with working examiners, pilot a small cohort, plan dedicated training time, and let usefulness drive voluntary use.Build them with the staff, start with a small group, set aside training time, and let usefulness — not a mandate — drive use.
An inaccurate citation or date is inherited by the examiner and reaches the taxpayer.A wrong source or date slips through to the taxpayer. Treat accuracy as a release gate: every output is pin-cited, a defined validation standard and Counsel / Practice Area sign-off precede shipping, and corrections are tracked.Don't ship a tool until its accuracy is checked and signed off; every answer shows its source, and fixes are tracked.
The redaction core misses something and exposes sensitive detail — the single highest-blast-radius failure.The tool that hides sensitive details misses one. Treat redaction as assistive and human-verified — it suggests; a person clears every page; it never marks content safe on its own.It only suggests — a person checks and clears every page by hand. It can never mark a page safe on its own.
Tools tuned on synthetic data encode wrong assumptions about messy real records (scanning noise, missing years, malformed forms).Tools built on clean made-up data stumble on messy real records. Run an explicit synthetic-to-real validation step before any tool leaves the pilot, and state the gap openly rather than hide it.Test deliberately against real-world mess before any tool leaves practice mode, and say so openly.
Over-reliance lets a tool's output read as a determination rather than a reviewed indication.People start treating a tool's answer as the decision. Embed indication-not-determination, read-only, post-selection, and human-in-the-loop in every output, with a review log on each result.Every result is clearly marked as something to review, never the decision, with a person in the loop and a record kept.
Over-reliance on a single contractor for the environment, AI assist, or plumbing is a single point of failure on the critical path — and no one is set up to maintain the tools after Day 90.Too much rides on one outside vendor, and no one owns upkeep after 90 days. Anchor on field-built, agency-owned tooling; the IRS receives all source and AI-assisted code with clear ownership and licensing, and a named maintenance owner is in place before the pilot ends.Keep the work and the code in government hands with clear ownership, and name who maintains it before the pilot ends.
Before week one

Day 0 — The Enabling Decision

The first decision of the pilot is not what we build, but where it can legitimately be built. This is a software-delivery effort, and software delivery runs on a standard toolchain — version control, issue tracking, and build/deploy automation, alongside AI-assisted development. A standard compliance-issue laptop is correctly configured for casework, not for building and shipping software — a known provisioning constraint, not a gap in IT. All three paths share the same guardrails: synthetic or sanitized data only, no Federal Tax Information ever in the environment, and the IRS owning every deliverable.The first decision isn't what we build — it's where we're allowed to build it. Building software needs proper software tools (a place to manage the code, track the work, and test it), and a standard casework laptop isn't set up for that — by design, and that's fine. Whatever we use: made-up data only, no real taxpayer information, and the government owns everything we build.

Government-Provisioned Modernization Device
Treasury or the modernization program issues a device pre-configured with an approved AI-assisted development environment, keeping all work on government hardware end to end. Highest sovereignty; longest stand-up time.Set up a fully approved government workspace from day one. Safest, keeps everything on government equipment; takes the longest to stand up.
Short-Term Detail to Vetted Partner
A scoped 90–120 day detail places the work inside a vetted private partner's secure, accredited development environment, with all finished tools returned to and owned by the IRS. A familiar federal mechanism; pace depends on vetting and agreement turnaround.Temporarily work inside a vetted partner's approved, secure space, handing all finished tools back to the IRS. A familiar arrangement; pace depends on the paperwork.
Authorized Use of an Equipped Environment
The IRS authorizes use of an already-provisioned, already-tooled secure environment for synthetic-data-only prototyping, with every deliverable handed to and owned by the IRS. Fastest to a first working tool because the setup cost is already paid.Get permission to use an already-set-up secure workspace for made-up-data practice only, with everything handed to the IRS. Fastest to a first working tool, because the setup is already done.
Across all three paths the rules are identical and non-negotiable: synthetic or sanitized data only with no FTI ever in the environment, the IRS owning and receiving every deliverable, and the arrangement temporary and bounded to the pilot window.Same rules on every path: made-up data only, no real taxpayer information ever, the government owns and receives everything, and the arrangement is temporary — just for the pilot.
A second Day-0 enabler: approved, direct access to a small group of front-line employees in each division throughout development — they are the source of tools that get used. Broader stakeholder coordination is sequenced to the pre-launch gate, not the build.The second thing to approve up front: direct access to a few frontline staff in each area throughout the build — they're the reason tools turn out useful. The wider coordination comes before launch, not during.
Three things to settle in week one, each with a named owner — these are the real timeline-killers: the workspace's security authorization, the synthetic-data and privacy sign-off, and Counsel/Disclosure review of the redaction approach. Surface them as gates, not deferred coordination.Three things must be nailed down in the first week, each with someone responsible — these usually cause the delays: permission to use the workspace, the OK to use made-up data, and a legal review of the tool that hides sensitive details.

12-week roadmap

Editable — rename, reschedule, add or remove items, and check them off.
A living document — deliverables and sequencing are proposed starting points and will be refined with leadership input. Anything here can be edited live in this meeting: drag the handle (or use the ▲▼ arrows) to reorder and set priority, rename a deliverable, change its target week, add or remove items, and check off what's done.
Weeks 1–4
Foundation — highest-value, hand-done work first
Weeks 5–8
Expansion — issue development & analysis
Weeks 9–12
Advanced — specialized & cross-cutting
Commitment: Core committed · Target planned · Stretch if time allows Under-promise, over-deliver — we commit to the core and treat the hard ones as stretch. Click a tag to change it.

Shared foundation

Built once, reused across all three tracks.

Six tools per track, one common spine. Each tool is a thin, governed layer on a shared base — so output stays consistent, every tool de-duplicates against the system of record, and the marginal cost of the next tool falls.Six tools per area, but one shared base underneath. Each tool is a thin layer on top, so they stay consistent and the next one costs less to build.

Read-Only Access Layer
Single ingest-only connection to ECM, IDRS, AIMS, and RGS within the authorized boundary; reads from each and writes back to none.One read-only connection to the official case systems — it can look, but never change anything.
Authority & Citation Engine
Shared IRM, IRC, and CFR lookup so every indication carries a verifiable pin-cite the examiner can check against source.A shared lookup so every answer comes with its source, which the staff member can check.
Indication Framework
Standard output contract marking each result an examiner-reviewed indication, never a determination, with disclaimer and review log.A standard format that labels every result as something to review, never a final decision, with a record of what happened.
PII/FTI & Parsing Core
One extraction and redaction core for bank, 433-A/B, M-3, and 1118/1099 documents, enforcing Pub 1075 data handling. Redaction is assistive and human-verified — it suggests; a person clears every page.One shared piece that reads documents and hides sensitive details, built to handle protected information safely. It assists — a person always confirms before anything is shared.
Workbench Shell & Export
Common examiner interface with export to IDR, workpaper, and PDF formats, shared so every tool looks and behaves consistently.A common screen and export (to the usual letter, workpaper, and PDF formats) so every tool looks and works the same.
Audit & Post-Selection Guard
Usage logging and feedback capture feeding the success measures, plus a guardrail confirming tools never touch selection, scoring, or targeting.A running record of use and feedback, plus a built-in rule that these tools never touch who gets selected or scored for audit.

Measuring success

Leading indicators — what we will measure in the first 90 days.

In ninety days we can credibly measure delivery, adoption, and output integrity. Downstream effects on cycle time and case quality are the hypothesis these indicators are designed to test — baselines and targets set with the field after the pilot, never pre-promised.What we'll actually measure in the first 90 days — early signals, not promised outcomes. The bigger payoff is what these are testing toward, set with the field after the pilot.

Deliverables shipped & validatedHow many tools we finished that passed the accuracy bar and legal review
Deliverables shipped & validated — tools delivered against the planned slate that clear the accuracy standard and Counsel review, not raw build count.How many tools we finished that passed the accuracy bar and legal review — quality-cleared, not just built.
Front-line task completed end-to-endSomething you can point to: a set number of frontline staff finish a whole practice task using a tool, each with a full record of sources and what happened
Front-line task completed end-to-end — a concrete artifact: a target number of front-line staff complete a real (sample-based) exam or collection task using a tool, each with a logged source trail and audit record.Something you can point to: a set number of frontline staff finish a whole practice task using a tool, each with a full record of sources and what happened.
Pilot-cohort adoptionHow many staff actually use the tools, and keep using them
Pilot-cohort adoption — active users and repeat use per office in the pilot, as a signal of voluntary pull rather than mandated activity.How many staff actually use the tools, and keep using them — a sign they're genuinely helpful.
Citation-verification rateHow often a tool's answer matches its stated source when checked
Citation-verification rate — the share of outputs whose underlying IRM, IRC, or CFR reference checks out; our core output-integrity signal.How often a tool's answer matches its stated source when checked — our main accuracy signal.
Examiner-rated usefulness & trustWhat staff say in quick surveys about how useful and trustworthy the tools are
Examiner-rated usefulness & trust — short pulse-survey ratings of usefulness and confidence so the field's own judgment guides what continues.What staff say in quick surveys about how useful and trustworthy the tools are.
Reviewer-flagged correction rateHow many fixes reviewers flag per tool, so we catch problems before any wider use
Reviewer-flagged correction rate — defects and corrections reviewers flag per tool, surfacing where a tool needs rework before scale.How many fixes reviewers flag per tool, so we catch problems before any wider use.
This is a working roadmap — deliverables and timelines are proposed starting points and will be refined with leadership input. Not yet an official IRS commitment, schedule, or system. · ★ America 250