Services AI Integration

AI where it earns its place.

Most AI budget buys a demonstration. This is about the other thing — the work your week is actually made of: the same documents, the same exports, the same eleven questions answered after hours. We find the part of it that repeats, take it apart, and hand back something that runs. Private where the material needs it, measured everywhere, and documented so it still works after we leave.

Partslocal LLMs · retrieval · Whisper · image models · Python · APIs

A finger on the power button of a small server box with its status light lit, a laptop beside it showing a private dashboard; hands typing at a screen of assistant replies; and a voice waveform on a phone held beside the text it was transcribed into.

§01 What it looks like in practice

Four shapes this work takes.

Almost every job we have built lands in one of these four. The names matter less than the test they share: something repeats, a person is doing the mechanical part of it, and the result can be checked by a machine before a person signs it off.

Workflow automation

The task done forty times a month, done the same way every time. Recurring work eats creative hours a few minutes at a time, and hand-made output drifts — version ten never quite matches version one.

PartsPython · APIs · git

Document & media pipelines

One source, rendered out to every format it is asked for — print-grade PDF, editable Word, a deck — and never retyped into any of them. Long video and audio chunked, processed and reassembled unattended.

Partsffmpeg · headless Chrome · LibreOffice · Python

Private assistants

Answers drawn from your own manuals and documents rather than from a model's general knowledge — run on your hardware where the material is sensitive, with a handoff to a person built in from the start rather than bolted on after the first bad answer.

Partslocal LLMs · retrieval · APIs

Generation at volume

Thirty of something, in one style, from one source of truth: artwork, narration, translation, worksheets, review material. Graded against a rubric before anything downstream is built from it, and finished by hand where hands still show.

Partsimage models · Whisper · text-to-speech · Markdown

§02 How the work goes

Five stages, and no surprises in the middle.

The same five stages whether the thing is a dubbing pipeline, an assistant over a shelf of manuals, or a document machine that turns one file into six. Each stage ends in something you can hold — a list, a working pilot, a measurement, a runbook — so the project can be stopped at any of them without leaving you with nothing.

01 · AUDIT Find theclicking 02 · PILOT One task,proven 03 · BUILD Madeboring 04 · MEASURE Checked,not assumed 05 · HAND OVER Yours torun
Swipe across — Every stage hands over an artefact, not a status update. The one that gets skipped in this industry is 04, which is why so much of this work is impressive in a meeting and quietly abandoned by the following quarter.

01 · Audit

Find the clicking

We follow the work, not the org chart: what repeats, how often, and how much of it is a person moving data between two windows.

What happens

Every candidate task gets three numbers — how many times a month, how many minutes each time, and how much real judgement it needs. Anything low on judgement and high on repetition is a candidate. Anything that turns on a human decision stays with the human, and gets marked that way in writing so nobody comes back to it in six months with the wrong idea.

The honest finding is often that the task irritating everyone is not the task costing anything. Both go on the list, with the hours attached, and you decide which one is worth the money.

What you get

  • A ranked list of candidates with hours attached to each
  • A column for the work we would refuse to automate, and why
  • One recommendation for the pilot, with the reason it goes first

02 · Pilot

One task, proven

One job, end to end, on your real material — never a demonstration on sample data, because sample data is where every AI project looks good.

What happens

A pilot exists to answer two questions that no proposal can: what does one run actually cost in time and money, and is the output good enough that nobody quietly redoes it by hand afterwards. Both get numbers.

“Good enough” is given a number before anything is built. On our own dubbing pipeline that number was about 93% of frontier quality — not perfection, but nothing that makes a listener wince. Writing it down early is what stops the review round becoming an argument about taste, and it is the difference between a project that ships and one that is still being polished at Christmas.

What you get

  • A working pilot, running on your own material
  • Measured cost and time per run, against what it costs you today
  • A go or no-go you can defend to somebody holding the budget

03 · Build

Made boring

The pilot becomes something dull and dependable: resumable, logged, version-pinned, and legible to the next person who opens it.

What happens

Resumable matters more than it sounds. Fixing one bad line in a two-hour job should cost you that line, not the two hours — so the work is chunked, checkpointed, and safe to re-run from where it broke. Model versions are pinned, because an unpinned pipeline is one that changes its output on a Friday without telling anyone. Where the material is sensitive, it runs on hardware you control; where it is not, the cheapest capable model wins.

It is also written to be read. Naming, comments and structure are part of the deliverable, on the assumption that the person maintaining it is not us.

What you get

  • The pipeline, running unattended, with logs that say what it did
  • The source — yours, versioned, no black box in the middle
  • Pinned versions, so an upstream update cannot quietly change your output

04 · Measure

Checked, not assumed

Output is checked by something that is not the thing that produced it — and then by a person, whose approval is never quietly overwritten.

What happens

The failure mode of this technology is not gibberish; it is fluent, confident and wrong, and it survives every short demonstration. So the check has to be mechanical. In the dubbing pipeline, every generated line of speech is transcribed back and compared against the text it was supposed to say — a machine listening to a machine, because a human spot-check of ten seconds proves nothing about the other fifty-nine minutes. In content work it is a grading pass against a rubric, run before anything downstream is generated from the master.

Then a person reads it. And a line a person has approved is never regenerated behind their back — reviewed work is kept, which sounds obvious until you meet a pipeline that does not do it.

What you get

  • A quality gate that runs on every job, not on the good days
  • A record of what was checked and what it scored
  • A number you can put in front of a board instead of an adjective

05 · Hand over

Yours to run

A workflow nobody on your team can run is a workflow you are renting. The handover is the job, not the paperwork after it.

What happens

Training is a working session on your own material rather than a slide deck: how to run it, how to change the parts you will certainly want to change, how to tell when it has gone wrong, and exactly what to do then. The written guide is aimed at somebody reading it at two in the morning with no one to ask — the same standard as the recovery guides we write for networks.

After that it runs without us. No retainer, no lock-in, no gatekeeper: if you come back for the next thing, it should be because the last one worked.

What you get

  • A training session with the people who will actually run it
  • A written runbook and a recovery guide, plain enough to follow cold
  • Everything in your hands — source, prompts, settings, documentation

§03 The toolkit

Parts, and why each one is there.

Nouns, not adjectives. These are the pieces most of this work is assembled from — chosen per job rather than per fashion, and most of them run on a machine you already own.

PartWhy it is there
Local LLMsThe model runs on your hardware, so the documents never leave the building to be read. Smaller and slower than the frontier, and for most document work you cannot tell the difference. Run through LM Studio or Ollama on Apple Silicon or a Linux box with a GPU.
Frontier APIsFor the jobs that genuinely need the strongest reasoning available and involve material you would be comfortable emailing. Priced per token, which is why the pilot measures the cost of a run before anyone commits to ten thousand of them.
RetrievalAnswers pulled from your own manuals, contracts and notes rather than from the model's general knowledge. This is what turns “a chatbot” into “our chatbot”, and it is what makes “I don't know, here is a person” a supported answer rather than an invented one.
WhisperSpeech to text with timestamps, locally, fast on Apple Silicon through MLX. Used forwards for transcription and subtitles — and backwards as the checker, transcribing generated audio to confirm it said what it was told to.
DemucsSeparates a voice from the music and the room behind it, so a dub or a clean-up does not take the soundtrack with it.
Voice cloningF5-TTS and its relatives speak each line in a voice cloned from a reference sample of twelve seconds or less. Curate one reference per speaker up front — we learned that the expensive way.
Image modelsThe current generation gets artwork to roughly eighty per cent. The last twenty is still a person with a pen, and the difference is visible from across a room.
PythonOne core driving the whole chain, with a resume flag, so a fix costs you the broken step rather than the whole run.
ffmpegCuts long media into workable chunks at silent gaps, then stitches the processed parts back together without a seam.
Headless ChromeRenders one HTML source into print-grade PDF. With LibreOffice alongside for editable Word output, it means a document is written once and never retyped into a second format.
gitEvery pipeline is versioned. A bad run can be walked back to the exact state that produced it, which is the difference between a bug and a mystery.

NoteNo framework religion. If a job is better served by something not on this list, it gets used and the reason is written down.

§04 Where it runs

Local, cloud, or both.

sensitive local model ordinary an API a person your hardware priced per token always
Both paths end in the same place. The person at the end is not a formality — it is the step that makes the rest defensible.

The split is a design decision, not a religion. Three questions settle it, and they are worth asking before anyone shows you a product:

  1. Would you email this to a stranger? If the answer is no — patient records, employment files, contracts, anything under an agreement you signed — it runs locally, on hardware you control, and the question of whose servers saw it never arises.
  2. Does the task need the best reasoning in the world, or competent reasoning? Most of it needs competent. Summarising, extracting, reformatting, translating, tagging and answering from a manual are all well inside what a local model does on a laptop.
  3. Does it run twice a week or ten thousand times a month? Per-token pricing is cheap in a pilot and occasionally startling at volume. The measurement from stage 02 is what makes that a forecast instead of a surprise.

Most businesses end up with both, wired into one pipeline: the private half local, the heavy reasoning on an API, and a person at the end of it either way. Which parts sit where is written down and handed over with everything else, so the answer to “where does our data go?” is a diagram rather than a shrug.

  • Runs on hardware you own
  • Source is yours
  • Documented for 2 a.m.
  • No retainer

§05 Not a sales page

How this usually goes wrong.

Six ways we have watched AI projects fail, including our own. If you are talking to somebody else about this work, these are fair questions to put to them.

  1. Fluent garbage. The output reads beautifully and is wrong. It is the one failure that survives a demonstration, because a demonstration is short and the prose is excellent. The fix is mechanical checking by something other than the generator, plus a person who reads — not a longer demo.
  2. Automating the wrong task. The job that annoys everybody is rarely the job that costs anything. Without the audit you spend the budget on the loudest complaint and the week looks exactly the same afterwards.
  3. No number for “good enough”. Without one agreed in advance, every review becomes a debate about taste, the goalposts move each time somebody new looks at it, and the thing never ships.
  4. The pilot that never became a habit. It worked, it impressed the room, and three months later everyone is back to doing it by hand — because nobody was trained, nothing was written down, and the one person who understood it got busy.
  5. An upstream update that changed your output. Unpinned versions mean the model underneath you can be replaced without notice. Everything still runs; the results are subtly different; nobody notices for a fortnight.
  6. Handing over judgement. These tools are good in the middle of a task and poor at both ends — deciding what the task is, and deciding whether the answer is right. Those two ends belong to people, and a system designed as though they don't will eventually make a decision nobody chose.

§06 Straight answers

The questions we actually get.

Will our data be used to train somebody's model?

Not if it runs locally — the material never leaves hardware you control, which is the whole reason local models are on the list. Where an API genuinely earns its place we use one whose terms say business data is not trained on, and you are told which steps touch it before anything is sent. If the answer has to be that nothing leaves the building, that is a constraint we design to rather than a compromise we ask you for.

Do we need to buy new hardware?

Usually not. Most of the local work in this studio runs on the Mac that was already on the desk — a current Apple Silicon machine with enough memory is a capable AI workstation, and our dubbing pipeline was built on one. If a job genuinely needs more, the pilot says so before you spend anything.

What does a first project look like?

One task, carried from audit to handover — weeks rather than quarters. Starting small is not caution. The first finished thing teaches you more about the next five than any amount of planning, and it does it while the budget is still small enough to change your mind.

What happens when it gets something wrong?

It will, and the design assumes so. A wrong answer should be visible rather than silent: a checker flags it, the log says which step produced it, and not knowing is a supported answer with a person on the other end of it. The systems worth worrying about are the ones that have never been seen to be wrong.

Is somebody really available at three in the morning?

The assistant is. A person is not, and a system built properly says so rather than implying otherwise. It answers what your own documents answer, it hands off what they do not, and the handoff is timestamped so nobody is left wondering whether a reply is coming. Round the clock means the repeat questions stop waiting for office hours — not that somebody is awake for them.

Does this replace staff?

Not in the work we take. What gets automated is what nobody wanted: retyping, re-exporting, chasing one file through four applications. The measure of a good result is a week with more judgement in it and less clicking, done by the same people.

What do you need from us to start?

The real material for one task — real files, real edge cases, the ugly ones especially — and an hour with the person who actually does the job rather than the description of the job. Sample data is where every AI project looks good.

Who owns what you build?

You do. The source is yours, the runbook is yours, and it runs without us. No retainer and no lock-in: if you want us for the next thing, it should be because the last one worked.

Something that repeats every week?

Describe the week rather than the technology. You will get an honest answer about what is worth automating — including the parts that are not.

Leave a note See the AI work

Alsothe dubbing pipeline, taken apart · thirty lessons, one command · workflows that give the hours back · networks & IT