← Back to Blog
26 August 2026
My journey to analytic obesity
I am crestfallen to admit that only this week I learned what a tax data strategy actually is. I'd been saying the phrase for years without anyone ever asking me to define it, which is how most of us get away with most things. I was hoping it would turn out to be something clever, something with weight to it, the New Testament crossed with a Latin to Irish translation handbook and a C++ training manual. It is little more than a checklist, I have been, to put it mildly, vexed since.
A tax data strategy is knowing where your information actually comes from, what it is, how it's used, and what you already have versus what you have to go and create, and that's the whole thing, nothing more dressed up than that, just an answer that makes you feel a bit stupid for not having said it yourself before.
Knowing where a number lives isn't the same as having it, though, and that's the part that gets skipped. The next step is pulling that number the same way every single time, straight from where it's born, into one structured extract that everything downstream reads from, not five different exports stitched together by hand at month end. One feed, one format, one place a compliance process, or a model, can actually query instead of a fresh spreadsheet rebuilt from nothing every quarter. That's the dataset. It sounds unglamorous because it is, but it's the difference between a data strategy and a nice conversation about one.
It doesn't stop at the ERP either, which makes sense, it is answers to simple questions like: Do you record calls in Zoom or Teams, and do you know that choice decides whether a transcript of an internal conversation now exists, searchable, sitting somewhere with its own retention clock running that nobody's watching, and who transcribes the Teams call, and where does it land, and did anyone actually agree it should exist in the first place. `when do I need to let someone know I built an agent, I created a number that will be used on a financial transaction, or when I get it to summarise a memo from a firm? None of that is an IT question, it only dresses up as one.
None of it gets solved by writing it down once and hoping, either. It gets solved by someone actually owning each of those answers, in writing, and that someone being a named person, not a shared inbox, because a data strategy without an owner is just a very expensive opinion.
The line that stuck with me most though was: stop pointing AI at repetitive, rules-based work, because that's an RPA job, not an AI one, and confusing the two is where most of this goes wrong. Think of it this way, you have a 7 hour flight to New York with a colleague, the aim of the game is to pass the time and arrive while still sane. An RPA is me, watching Jurassic Park and then First Blood, in that order, every single time, never bored, never improvising, landing exactly where it's told to land. AI is my colleague on the same flight, sharp as anything, who spends the first forty minutes deciding between three films, changes seat once, and ends up deep in conversation with the stranger next to them about religion or politics when they just want to watch Universal Soldier, and still arrives safe and sane, but having just done something far more interesting with his time than me. Both of them get you there. One of them is built to think, and every time you ask it to do what a bot would do for free, you're wasting the thinking and paying for the privilege.
So before you build a workflow on top of any of this, or let a model anywhere near it, ask four questions, in order, and don't skip ahead. Where does this number come from, really, not where it's copied to, but where it's born. What is it, in plain terms, not the field name in the ERP but what a human would call it if you made them say it out loud. How is it used, and by whom, and does that person or system actually need it, or have they just always had it. And last, what do you already hold versus what are you manufacturing fresh every quarter, because nobody wrote down that it already exists somewhere upstream. Skip any one of those and you're not saving time, you're just moving the risk further downstream and calling it automation, because a model built on data you can't trace back to its source will hand you a wrong answer with exactly the same confidence as a right one, and it'll do it faster than Excel ever could.
Take one return, one line on it, and walk it backwards through those four questions before you touch a piece of software or a prompt box. Nine times out of ten the number is born in a system you already own, gets exported into a spreadsheet by someone who was never once asked why, gets used by three people who each think a different one of them owns it, and half of what you're rebuilding every quarter already exists two systems upstream, quietly, under a name nobody thought to tell you about. Do that exercise properly, on one filing, before you automate a single step of it.
What I actually want is to get to the point where I can type a sentence in plain English and have a compliance process run without touching Excel, without me pulling an SAP report, without me doing anything at all except applying tax judgement at the one point judgement is actually needed. I want to be Homer Simpson washing himself with the rag on a stick, so entirely removed from the mechanics of the thing that the only job left to me is a bit of light commentary, remarking that input credits are up in France again while somewhere behind me the actual washing takes care of itself.
None of it works without the sourcing. Until then, I'm no Homer, I'm just a fella with a good line about France.
Never miss a tax update
Get the week's most relevant tax news delivered every Friday. Free, no ads, no data selling.
Subscribe to the Newsletter