← Back to Blog
9 July 2026
I spent 60 hours making a tool 5% better.
A note before we start. I recently built a tool that runs a recurring tax risk review across a product portfolio. Out of respect for my employer, I am not going to describe what it covers, what it concludes, or anything about the underlying business. What I can share, and what I think is actually useful to other in-house tax people, is how I built it, where the time went, and what I got wrong. The tax content is deliberately generic, but lessons are genuine.
A lot of in-house tax work is the periodic review of a large number of items against a framework: products against taxability rules, transactions against place of supply logic, business models against whatever regime is evolving this year. The traditional answer is an in-depth external review, which costs tens of thousands, takes months, and is out of date as soon as it is converted to .pdf, because products, pricing models and rules all keep moving. Worse, by the time anything shows up in the financial data, the interesting conversation happened eighteen months earlier in a strategy deck. I wanted a monthly control instead of a periodic memo.
The build was less impressive than that sounds. I started by writing down the tax logic in plain English: which characteristics actually matter for the analysis and, just as importantly, which do not. No tooling, no data, just the framework a good reviewer would apply, written out so that it could be applied consistently by something that is not me. I then sent that document to external advisers and asked them to pull it apart, which they did, thoroughly. Several of my assumptions did not survive. That was mildly bruising and entirely the point. The revised logic became a locked assumptions log, and everything else sits on top of it.
The stack is not much of a stack. A periodic revenue extract, a mapping table to translate internal system product codes into something a human can read, a rules register, and an LLM that reviews each item in the catalogue individually (from our homepage, do it does not need upkeep) against the framework, with a stated reason for every flag it raises. Each run validates the data first, because an analysis built on a broken extract is worse than none, then classifies, applies the rules, and quantifies. Items already handled compliantly come out as reconciliation points rather than headline risks, because nothing erodes trust faster than presenting a solved problem as an exposure. The output is the same every month: a short report for leadership, a workbook, the assumptions log and an action list. Same structure every run, so movement is visible month on month.
The first version was honestly a bit rubbish. One extract, the philosophy document, and the model classifying items, with none of the supporting machinery. It took me somewhere between 15 and 20 hours to build, and it gave me some genuinely useful insights, particularly on what could be ruled out, but the calibration was off. My first framework was too generous and cheerfully flagged most of the portfolio as sensitive, which would have made for a very short meeting with the boss. Tightening it risked missing the things that were the actual point of the exercise. The fix was giving up on a single yes-or-no flag and classifying into explicit categories instead, running from clearly sensitive through grey-zone down to an honest "no idea, ask the product owner".
Then I spent another 40 to 60 hours rectifying, refining and polishing, and here is the uncomfortable arithmetic: that second investment, nearly four times the size of the first, did not produce four times the improvement. It did not produce anything close. The rough first draft got the tool to about 80% of its usefulness in a fraction of the time, and the long tail of refinement mostly bought me tidier edge cases and a nicer assumptions log. If I did it again I would stop earlier. The 80% version was already answering questions the expensive memo could not, and for a working control that is probably the right place to declare victory and start running it. Blank page to a working monthly run within a couple of weeks, most of which, in hindsight, was me arguing with my own assumptions past the point of diminishing returns.
Was it clever? Not especially. And to be clear about what it is not: this is not a substitute for knowing tax, nor is anything it produces a substitute for research and judgement by a professional. If I could not have written the philosophy document, or recognised which of the advisers' challenges were fatal and which were noise, the tool would just be automating my ignorance at scale. What it actually does is let me flex the tax muscle far more often than the old model allowed. The judgement that used to get exercised once a year, at memo time, now gets exercised every month, against fresh data, in minutes rather than months.
The AI does the reading and the consistency, which are the parts humans are worst at, and the advisers now spend their time on challenge and exceptions instead of first-pass review, which is a better use of them and a more interesting brief. If your work involves periodically reviewing a large number of items against a stable framework, and most in-house indirect tax work does, the same approach will probably embarrass you into building something similar. Start with the plain-English framework, get it challenged by someone who will enjoy dismantling it, build the rough version in a couple of weekends, and stop at 80%.
Never miss a tax update
Get the week's most relevant tax news delivered every Friday. Free, no ads, no data selling.
Subscribe to the Newsletter