Workflow automation
The task done forty times a month, done the same way every time. Recurring work eats creative hours a few minutes at a time, and hand-made output drifts — version ten never quite matches version one.
PartsPython · APIs · git
Services AI Integration
Most AI budget buys a demonstration. This is about the other thing — the work your week is actually made of: the same documents, the same exports, the same eleven questions answered after hours. We find the part of it that repeats, take it apart, and hand back something that runs. Private where the material needs it, measured everywhere, and documented so it still works after we leave.
Partslocal LLMs · retrieval · Whisper · image models · Python · APIs
§01 What it looks like in practice
Almost every job we have built lands in one of these four. The names matter less than the test they share: something repeats, a person is doing the mechanical part of it, and the result can be checked by a machine before a person signs it off.
The task done forty times a month, done the same way every time. Recurring work eats creative hours a few minutes at a time, and hand-made output drifts — version ten never quite matches version one.
PartsPython · APIs · git
One source, rendered out to every format it is asked for — print-grade PDF, editable Word, a deck — and never retyped into any of them. Long video and audio chunked, processed and reassembled unattended.
Partsffmpeg · headless Chrome · LibreOffice · Python
Answers drawn from your own manuals and documents rather than from a model's general knowledge — run on your hardware where the material is sensitive, with a handoff to a person built in from the start rather than bolted on after the first bad answer.
Partslocal LLMs · retrieval · APIs
Thirty of something, in one style, from one source of truth: artwork, narration, translation, worksheets, review material. Graded against a rubric before anything downstream is built from it, and finished by hand where hands still show.
Partsimage models · Whisper · text-to-speech · Markdown
§02 How the work goes
The same five stages whether the thing is a dubbing pipeline, an assistant over a shelf of manuals, or a document machine that turns one file into six. Each stage ends in something you can hold — a list, a working pilot, a measurement, a runbook — so the project can be stopped at any of them without leaving you with nothing.
01 · Audit
We follow the work, not the org chart: what repeats, how often, and how much of it is a person moving data between two windows.
Every candidate task gets three numbers — how many times a month, how many minutes each time, and how much real judgement it needs. Anything low on judgement and high on repetition is a candidate. Anything that turns on a human decision stays with the human, and gets marked that way in writing so nobody comes back to it in six months with the wrong idea.
The honest finding is often that the task irritating everyone is not the task costing anything. Both go on the list, with the hours attached, and you decide which one is worth the money.
02 · Pilot
One job, end to end, on your real material — never a demonstration on sample data, because sample data is where every AI project looks good.
A pilot exists to answer two questions that no proposal can: what does one run actually cost in time and money, and is the output good enough that nobody quietly redoes it by hand afterwards. Both get numbers.
“Good enough” is given a number before anything is built. On our own dubbing pipeline that number was about 93% of frontier quality — not perfection, but nothing that makes a listener wince. Writing it down early is what stops the review round becoming an argument about taste, and it is the difference between a project that ships and one that is still being polished at Christmas.
03 · Build
The pilot becomes something dull and dependable: resumable, logged, version-pinned, and legible to the next person who opens it.
Resumable matters more than it sounds. Fixing one bad line in a two-hour job should cost you that line, not the two hours — so the work is chunked, checkpointed, and safe to re-run from where it broke. Model versions are pinned, because an unpinned pipeline is one that changes its output on a Friday without telling anyone. Where the material is sensitive, it runs on hardware you control; where it is not, the cheapest capable model wins.
It is also written to be read. Naming, comments and structure are part of the deliverable, on the assumption that the person maintaining it is not us.
04 · Measure
Output is checked by something that is not the thing that produced it — and then by a person, whose approval is never quietly overwritten.
The failure mode of this technology is not gibberish; it is fluent, confident and wrong, and it survives every short demonstration. So the check has to be mechanical. In the dubbing pipeline, every generated line of speech is transcribed back and compared against the text it was supposed to say — a machine listening to a machine, because a human spot-check of ten seconds proves nothing about the other fifty-nine minutes. In content work it is a grading pass against a rubric, run before anything downstream is generated from the master.
Then a person reads it. And a line a person has approved is never regenerated behind their back — reviewed work is kept, which sounds obvious until you meet a pipeline that does not do it.
05 · Hand over
A workflow nobody on your team can run is a workflow you are renting. The handover is the job, not the paperwork after it.
Training is a working session on your own material rather than a slide deck: how to run it, how to change the parts you will certainly want to change, how to tell when it has gone wrong, and exactly what to do then. The written guide is aimed at somebody reading it at two in the morning with no one to ask — the same standard as the recovery guides we write for networks.
After that it runs without us. No retainer, no lock-in, no gatekeeper: if you come back for the next thing, it should be because the last one worked.
§03 The toolkit
Nouns, not adjectives. These are the pieces most of this work is assembled from — chosen per job rather than per fashion, and most of them run on a machine you already own.
| Part | Why it is there |
|---|---|
| Local LLMs | The model runs on your hardware, so the documents never leave the building to be read. Smaller and slower than the frontier, and for most document work you cannot tell the difference. Run through LM Studio or Ollama on Apple Silicon or a Linux box with a GPU. |
| Frontier APIs | For the jobs that genuinely need the strongest reasoning available and involve material you would be comfortable emailing. Priced per token, which is why the pilot measures the cost of a run before anyone commits to ten thousand of them. |
| Retrieval | Answers pulled from your own manuals, contracts and notes rather than from the model's general knowledge. This is what turns “a chatbot” into “our chatbot”, and it is what makes “I don't know, here is a person” a supported answer rather than an invented one. |
| Whisper | Speech to text with timestamps, locally, fast on Apple Silicon through MLX. Used forwards for transcription and subtitles — and backwards as the checker, transcribing generated audio to confirm it said what it was told to. |
| Demucs | Separates a voice from the music and the room behind it, so a dub or a clean-up does not take the soundtrack with it. |
| Voice cloning | F5-TTS and its relatives speak each line in a voice cloned from a reference sample of twelve seconds or less. Curate one reference per speaker up front — we learned that the expensive way. |
| Image models | The current generation gets artwork to roughly eighty per cent. The last twenty is still a person with a pen, and the difference is visible from across a room. |
| Python | One core driving the whole chain, with a resume flag, so a fix costs you the broken step rather than the whole run. |
| ffmpeg | Cuts long media into workable chunks at silent gaps, then stitches the processed parts back together without a seam. |
| Headless Chrome | Renders one HTML source into print-grade PDF. With LibreOffice alongside for editable Word output, it means a document is written once and never retyped into a second format. |
| git | Every pipeline is versioned. A bad run can be walked back to the exact state that produced it, which is the difference between a bug and a mystery. |
NoteNo framework religion. If a job is better served by something not on this list, it gets used and the reason is written down.
§04 Where it runs
The split is a design decision, not a religion. Three questions settle it, and they are worth asking before anyone shows you a product:
Most businesses end up with both, wired into one pipeline: the private half local, the heavy reasoning on an API, and a person at the end of it either way. Which parts sit where is written down and handed over with everything else, so the answer to “where does our data go?” is a diagram rather than a shrug.
§05 Not a sales page
Six ways we have watched AI projects fail, including our own. If you are talking to somebody else about this work, these are fair questions to put to them.
§06 Straight answers
Not if it runs locally — the material never leaves hardware you control, which is the whole reason local models are on the list. Where an API genuinely earns its place we use one whose terms say business data is not trained on, and you are told which steps touch it before anything is sent. If the answer has to be that nothing leaves the building, that is a constraint we design to rather than a compromise we ask you for.
Usually not. Most of the local work in this studio runs on the Mac that was already on the desk — a current Apple Silicon machine with enough memory is a capable AI workstation, and our dubbing pipeline was built on one. If a job genuinely needs more, the pilot says so before you spend anything.
One task, carried from audit to handover — weeks rather than quarters. Starting small is not caution. The first finished thing teaches you more about the next five than any amount of planning, and it does it while the budget is still small enough to change your mind.
It will, and the design assumes so. A wrong answer should be visible rather than silent: a checker flags it, the log says which step produced it, and not knowing is a supported answer with a person on the other end of it. The systems worth worrying about are the ones that have never been seen to be wrong.
The assistant is. A person is not, and a system built properly says so rather than implying otherwise. It answers what your own documents answer, it hands off what they do not, and the handoff is timestamped so nobody is left wondering whether a reply is coming. Round the clock means the repeat questions stop waiting for office hours — not that somebody is awake for them.
Not in the work we take. What gets automated is what nobody wanted: retyping, re-exporting, chasing one file through four applications. The measure of a good result is a week with more judgement in it and less clicking, done by the same people.
The real material for one task — real files, real edge cases, the ugly ones especially — and an hour with the person who actually does the job rather than the description of the job. Sample data is where every AI project looks good.
You do. The source is yours, the runbook is yours, and it runs without us. No retainer and no lock-in: if you want us for the next thing, it should be because the last one worked.
Describe the week rather than the technology. You will get an honest answer about what is worth automating — including the parts that are not.
Alsothe dubbing pipeline, taken apart · thirty lessons, one command · workflows that give the hours back · networks & IT