I Took the Math Away From the LLM
How I rebuilt a monthly reporting pipeline so the AI writes sentences instead of inventing data, cutting monthly reporting time by 97%.
The task I was delegated
I write monthly wellness engagement reports for several multinational corporate clients, including 3 Fortune 500 manufacturing companies. These aren't internal memos nobody reads. They go straight to upper management, and they're part of the reason our service stays a line item instead of a chopping block target. The data has to hold up under scrutiny.
The workflow I inherited did not hold up to that standard. Prior to my onboarding, the raw Excel attendance export was fed directly into an LLM, asking it to generate pivot tables and a chart. Sounds efficient on paper, right? It wasn't.
LLMs are predictive text engines [1]. They are autocomplete on steroids, and they don't count rows in a spreadsheet. Ask one to aggregate 1000+ rows of attendance data and it will hand you numbers that were never there, with total confidence. It genuinely can't tell the difference between "correct" and "plausible." Either way, someone has to catch it before a client does.
That someone was a human (me), manually cross-checking every fabricated pivot table against the raw file before anything went near a client inbox. Ninety minutes per company. Three companies a month. And most of that ninety minutes wasn't spent working — it was spent hunting for a number the AI invented.
Separate the math from the writing
The old workflow asked one tool to do two different jobs: calculate exact numbers, and write about them in plain English. An LLM is decent at the second job. It has no business doing the first. So I split the pipeline in two, and made sure the two parts could never touch each other's work.
Part one — the math, done natively
I wrote a Google Apps Script [2] that reads every client's raw attendance sheet in the workbook, groups check-ins by class type — Open Gym, Yoga, Strong Nation, whatever's running that month — and calculates totals and per-class averages with plain spreadsheet logic. No AI involved anywhere in this step. It builds a dashboard sheet with a native pie chart for each company, and it runs across all three companies in one execution, in about five seconds.
Part two — the writing, constrained to a fact-checker's job
A Python script [3] reads the already-verified dashboard. Because the numbers are locked in by this point, there's nothing left to hallucinate; the script does two things. It renders a branded donut chart with Matplotlib [4]. And it sends only the pre-calculated totals to an LLM through the OpenRouter API [5], pinned down with a system prompt [6] that gives it exactly one job: write a two-to-three sentence Attendance Overview, professional tone, no fluff, no subject line, no greeting, no bullet points. The model never sees a single raw data point, effectively shutting the door to hallucination before the LLM even gets in the room.
The script stitches the chart, the verified tables, and the LLM's paragraph into a ready-to-paste HTML email body. I open it, copy it into the email template, and send.
Later — dropping Apps Script entirely
Once the pipeline was stable, I kept looking at it and asking whether every piece still earned its place. The Apps Script step wasn't wrong, but it was a relay: raw data → Sheets → Apps Script → dashboard tab → Python. So I moved the calculation into Python itself, reading the raw spreadsheet directly. It didn't save meaningful time — the whole run was already under a minute. What it removed was a dependency. See Workflow C below.
Workflow breakdown
Three stages of the same pipeline, iterating away from both hallucination and unnecessary tooling. Green is automated, white is manual, red is where the old system fell apart.
Workflow A — Legacy LLM method (per company, original)
Workflow B — Hybrid automation pipeline (interim, all 3 companies)
Workflow C — Single-script pipeline (current, all 3 companies)
What actually changed
| Workflow | Condition | Time (all 3 companies) | Manual steps | Error risk |
|---|---|---|---|---|
| A — Legacy LLM | Raw data fed to LLM | ~270 min | 3 + full manual audit | High (fabricated aggregations) |
| B — Hybrid pipeline | Math native, LLM writes only | ~8 min | 2 | Negligible |
| C — Single-script pipeline current | Math native, one tool | ~8 min | 2 | Negligible |
The bigger number isn't even the time. It's the zero. A report that used to need a full manual audit before it could leave the building now needs a glance. The math was never wrong to begin with, because the math was never the LLM's job. Workflow C doesn't move that needle further — it just means one fewer tool between the raw data and the finished report.
Contain your tools, not abandon them
The lesson here isn't "don't use AI for reports." It's "don't use AI for the part of the report where being wrong is a liability." Language generation and arithmetic are different disciplines that happen to share a chat window, and treating them as one job is how you end up with a beautifully formatted lie sent to a client's inbox.
Containment is the actual skill. Let deterministic code — Apps Script, Python — do the deterministic work: counting rows, summing columns, building a chart. Hand the LLM only verified, static facts and ask it to describe them. It never sees the raw data, so it can't invent what it never saw.
This scales further than one report. Anywhere a business workflow needs "read the numbers, tell me what happened," the same split applies: math stays in code, language stays with the model, and neither one is asked to cover for the other.
Workflow C taught a bonus lesson, after the hallucination problem was already solved. Folding Apps Script into Python didn't save meaningful time but it saved time on troubleshooting. By removing the dependency on two separate platforms, there's now one less thing to break. Sometimes, shaving off time in a stopwatch isn't the only measure of success.