Case Study · Operations & Data Integrity

I Took the Math Away From the LLM

How I rebuilt a monthly reporting pipeline so the AI writes sentences instead of inventing data, cutting monthly reporting time by 97%.

Role: Operations & Project Manager, Fit At Work Tools: Google Apps Script, Python, Matplotlib, OpenRouter API, Excel Time saved: ~262 min/month, recurring
C — Context

The task I was delegated

I write monthly wellness engagement reports for several multinational corporate clients, including 3 Fortune 500 manufacturing companies. These aren't internal memos nobody reads. They go straight to upper management, and they're part of the reason our service stays a line item instead of a chopping block target. The data has to hold up under scrutiny.

The workflow I inherited did not hold up to that standard. Prior to my onboarding, the raw Excel attendance export was fed directly into an LLM, asking it to generate pivot tables and a chart. Sounds efficient on paper, right? It wasn't.

LLMs are predictive text engines [1]. They are autocomplete on steroids, and they don't count rows in a spreadsheet. Ask one to aggregate 1000+ rows of attendance data and it will hand you numbers that were never there, with total confidence. It genuinely can't tell the difference between "correct" and "plausible." Either way, someone has to catch it before a client does.

That someone was a human (me), manually cross-checking every fabricated pivot table against the raw file before anything went near a client inbox. Ninety minutes per company. Three companies a month. And most of that ninety minutes wasn't spent working — it was spent hunting for a number the AI invented.


A — Action

Separate the math from the writing

The old workflow asked one tool to do two different jobs: calculate exact numbers, and write about them in plain English. An LLM is decent at the second job. It has no business doing the first. So I split the pipeline in two, and made sure the two parts could never touch each other's work.

Part one — the math, done natively

I wrote a Google Apps Script [2] that reads every client's raw attendance sheet in the workbook, groups check-ins by class type — Open Gym, Yoga, Strong Nation, whatever's running that month — and calculates totals and per-class averages with plain spreadsheet logic. No AI involved anywhere in this step. It builds a dashboard sheet with a native pie chart for each company, and it runs across all three companies in one execution, in about five seconds.

Part two — the writing, constrained to a fact-checker's job

A Python script [3] reads the already-verified dashboard. Because the numbers are locked in by this point, there's nothing left to hallucinate; the script does two things. It renders a branded donut chart with Matplotlib [4]. And it sends only the pre-calculated totals to an LLM through the OpenRouter API [5], pinned down with a system prompt [6] that gives it exactly one job: write a two-to-three sentence Attendance Overview, professional tone, no fluff, no subject line, no greeting, no bullet points. The model never sees a single raw data point, effectively shutting the door to hallucination before the LLM even gets in the room.

The script stitches the chart, the verified tables, and the LLM's paragraph into a ready-to-paste HTML email body. I open it, copy it into the email template, and send.

The old workflow asked the LLM to be a calculator and a copywriter at the same time. Splitting the two just means I am utilizing the LLM exactly as it was designed, to write about facts, not to invent them.

Later — dropping Apps Script entirely

Once the pipeline was stable, I kept looking at it and asking whether every piece still earned its place. The Apps Script step wasn't wrong, but it was a relay: raw data → Sheets → Apps Script → dashboard tab → Python. So I moved the calculation into Python itself, reading the raw spreadsheet directly. It didn't save meaningful time — the whole run was already under a minute. What it removed was a dependency. See Workflow C below.


Workflow breakdown

Three stages of the same pipeline, iterating away from both hallucination and unnecessary tooling. Green is automated, white is manual, red is where the old system fell apart.

Workflow A — Legacy LLM method (per company, original)

Download raw Excel attendance records
→
Feed entire raw dataset into LLM
→
Ask LLM to generate pivot table & chart
→
LLM hallucinates the aggregations
→
Manually audit every number against raw data
→
Manually draft email, risking date/subject errors
Total time: ~90 min / company · 3 companies/month = 270 min · Error-prone steps: 3

Workflow B — Hybrid automation pipeline (interim, all 3 companies)

Download raw Excel attendance records
→
Apps Script computes totals & averages, all companies at once
→
Python renders chart + fetches LLM overview from verified numbers
→
Script outputs ready-to-paste HTML email
→
Copy chart + HTML into email template, send
Total time: ~8 min / month · All 3 companies · Error-prone steps: 0

Workflow C — Single-script pipeline (current, all 3 companies)

Raw attendance data lands directly in spreadsheet
→
Python computes totals & averages natively, all companies at once
→
Same script renders chart + fetches LLM overview from verified numbers
→
Script outputs ready-to-paste HTML email
→
Copy chart + HTML into email template, send
Total time: ~8 min / month · All 3 companies · Error-prone steps: 0 · One tool instead of two

R — Results

What actually changed

97%
reduction in monthly reporting time (270 min → 8 min, across 3 companies)
0
hallucinated data points — the LLM never touches a raw number
34×
faster than the manual-audit workflow it replaced
Workflow Condition Time (all 3 companies) Manual steps Error risk
A — Legacy LLM Raw data fed to LLM ~270 min 3 + full manual audit High (fabricated aggregations)
B — Hybrid pipeline Math native, LLM writes only ~8 min 2 Negligible
C — Single-script pipeline current Math native, one tool ~8 min 2 Negligible

The bigger number isn't even the time. It's the zero. A report that used to need a full manual audit before it could leave the building now needs a glance. The math was never wrong to begin with, because the math was never the LLM's job. Workflow C doesn't move that needle further — it just means one fewer tool between the raw data and the finished report.


L — Learning

Contain your tools, not abandon them

The lesson here isn't "don't use AI for reports." It's "don't use AI for the part of the report where being wrong is a liability." Language generation and arithmetic are different disciplines that happen to share a chat window, and treating them as one job is how you end up with a beautifully formatted lie sent to a client's inbox.

Containment is the actual skill. Let deterministic code — Apps Script, Python — do the deterministic work: counting rows, summing columns, building a chart. Hand the LLM only verified, static facts and ask it to describe them. It never sees the raw data, so it can't invent what it never saw.

This scales further than one report. Anywhere a business workflow needs "read the numbers, tell me what happened," the same split applies: math stays in code, language stays with the model, and neither one is asked to cover for the other.

Workflow C taught a bonus lesson, after the hallucination problem was already solved. Folding Apps Script into Python didn't save meaningful time but it saved time on troubleshooting. By removing the dependency on two separate platforms, there's now one less thing to break. Sometimes, shaving off time in a stopwatch isn't the only measure of success.

Notes for non-technical readers

[1] Data hallucination is when an AI confidently invents information that was never in the source. It's not malicious — it's the model predicting what a plausible-sounding number looks like, rather than actually counting anything. ↩
[2] Google Apps Script is a scripting language built into Google Sheets, Docs, and other Workspace apps. It lets you automate tasks directly inside a spreadsheet without any AI involved — closer to a very capable macro than a chatbot. ↩
[3] A script is a small program that automates a repetitive task — in this case, reading verified numbers, drawing a chart, and asking an AI to write one paragraph about them. ↩
[4] Matplotlib is a Python library for drawing charts and graphs. It's used here to generate the donut chart that shows attendance split by class type. ↩
[5] OpenRouter is a service that lets a script talk to different AI models through one connection, instead of wiring up each AI provider separately. ↩
[6] A system prompt is the set of instructions given to an AI model before it sees the actual request — it sets the rules of the job. Here, it tells the model: write a short paragraph from these numbers, nothing else, don't touch the math. ↩