How local models run my personal podcast
September 2026
Every morning a ten minute news podcast shows up in a Discord channel on my phone. Two hosts go through the stuff I saved the day before (newsletters, articles, YouTube videos etc...) and then hand it over to a finance desk, a Knicks desk and a music desk. All of it gets made in my home office on one PC with a mid-range GPU in it, an RTX 4070 Super. That means the writing, the fact check, the voices and the mix. Nothing goes out to the cloud.
Listen while you read
A personal audio brief: how the show gets made
A short special made by the same pipeline this post describes: one host walks through the setup, then goes live to the music desk.
2 min 21 s · mp3 · made September 16, 2026
Why local
- The data stays home. I hand a local model every article I saved, my whole record collection and years of listening history. I would think twice about sending that pile to some service, and thinking twice is usually where I give up on a side project.
- Nobody is charging me by the token. With a cloud model every experiment has a little "is this worth it" tax on it. Once the card is paid for, running something fifty times costs electricity. That's how the voices got dialed in, a bunch of overnight bake-offs that only made sense because they were free.
- It keeps working. APIs change, prices change, models get retired, credits run out (mine did earlier this month and every job on the network quietly fell over to the expensive fallback). The model on my disk is still there tomorrow.
The big cloud models are still smarter and I use them all day for other stuff. Up until this week they were writing the script for this show too. As of today the script gets written on the card as well, so the whole thing is local (the cloud is still hooked up as a fallback in case the box goes down).
The hardware
- Off the shelf desktop build: Core i7, a bunch of RAM and a 12 GB RTX 4070 Super. It runs Linux and stays on.
- On the card: a 26 billion parameter Gemma model writes the script (mixture of experts, about 4 billion active, so it fits with room to spare). A 14 billion parameter Qwen model fact checks it. Whisper. And the text to speech model that does the voices. 12 GB is the number everything gets designed around.
- The collector: a little Mac mini with no graphics card. It gathers, schedules, hands the stories to the PC, waits, and copies the mp3 back.
How the show gets made
- Anything I save (articles, bookmarks, YouTube videos) lands in a database on the collector. At 7:15 an edition gets put together: a lead story, sections, running storylines grouped by plain text search. It's deterministic, so rerunning it is harmless.
- The Gemma model turns the top ten stories into a two host conversation, about 650 words. Numbers as words, acronyms spelled out, only things a person would actually say out loud. While it writes it has two tools: it can fetch a story's source page and it can search the web. A handful of calls at most, so the details on air come from the article itself.
- A second model from a different family reads it cold as a "doctor" and cuts anything the sources don't back up. It runs on the CPU on purpose so the card stays free. The whole writing stage takes about twelve minutes and most of that is the doctor thinking, which is fine at seven in the morning.
- The desks. Finance: a markets brief written at six and checked against its sources first. Knicks: the last 36 hours of coverage plus the next games, kept short. Music: below. Each desk is one block. The host sets it up, the correspondent answers in one go, the host moves on. That's the sound I wanted and it happens to be the cheapest thing to render.
- The script and two music clips get copied to the PC. A guard waits until 10 GB of video memory is free, then the voice model reads the show.
- Whisper transcribes the mp3 and diffs it against the script. Coverage per section gets posted with the show, plus the weakest chunk and a timestamp so I can spot check it. A chunk that fails badly gets re-read by a cloud voice, and a monthly rollup shows whether the drift is getting worse.
- The mp3 goes to Discord with the headlines and a copy lands in iCloud. If any step dies, a message in the channel says which one.
The voice trick
A 12 GB card is enough because of one trick. The voice model is Dia, 1.6 billion parameters, open weights, and it does five to twenty seconds of audio per call. So a ten minute show gets rendered in 50 word chunks and stitched, and the voices drift between chunks.
The fix: render one short anchor exchange with a fixed seed, once. Every chunk after that is generated as anchor plus new text, with the anchor audio as the prompt, then the anchor gets trimmed off the front. Each correspondent has its own anchor, so the Knicks guy always sounds like the Knicks guy.
The music desk
- Venue newsletters go to their own mailbox and get parsed into a concert database, one row per artist, venue and date.
- Every Sunday a taste model gets rebuilt from my Discogs collection, my Last.fm history and thumbs up or down on past picks.
- Every morning it ranks tonight's shows, pulls up to ten songs by tonight's artists into an Apple Music playlist (swapping out yesterday's), picks two for air and grabs the 30 second previews.
- On air the host introduces the song, the clip plays, and the concert segment says who's playing where tonight. A show can surface because of an opener I own three records by. On a slow night it skips the pick.
If you wanted to build one
A few things I would tell you if you were starting one of these:
- Keep the scripts dumb. The gathering and the sending are plain code and the model only writes.
- If a desk is missing one morning (no finance brief, venue email didn't come) the show should just run without it.
- Have a different model check the writer. Whisper checking the voice model has caught more problems than anything else IMO.
- Pick models by video memory first. 12 GB rules out a lot, and the mixture of experts ones are how you get around it.
- Watch what else lives on the card, a second model will evict the first one without a thought and gunk up the whole thing. Partly why the fact checker runs on the CPU.
- Never hurts to have a cheap cloud fallback model or provider plugged in. Local GPUs get unplugged by the news watchers sometimes.
If you're running something like this at home, curious what you've got it doing. Shoot me a note.
Jim