← jamescrews.com

How local models run my personal podcast

September 2026

Every morning a ten minute news podcast shows up in a Discord channel on my phone. Two hosts go through the stuff I saved the day before (newsletters, articles, YouTube videos etc...) and then hand it over to a finance desk, a Knicks desk and a music desk. All of it gets made in my home office on one PC with a mid-range GPU in it, an RTX 4070 Super. That means the writing, the fact check, the voices and the mix. Nothing goes out to the cloud.

Listen while you read

A personal audio brief: how the show gets made

A short special made by the same pipeline this post describes: one host walks through the setup, then goes live to the music desk.

2 min 21 s · mp3 · made September 16, 2026

Why local

The big cloud models are still smarter and I use them all day for other stuff. Up until this week they were writing the script for this show too. As of today the script gets written on the card as well, so the whole thing is local (the cloud is still hooked up as a fallback in case the box goes down).

Where the work happens My office What I save articles, links, video Collector gathers, schedules, ships the stories cloud fallback only PC with the 4070 Super writes the script voices, music, mastering fact-check on the CPU Whisper check mp3 to my phone
The blue boxes are hardware in my office. The dashed one only fires if the box is down.

The hardware

How the show gets made

The morning pipeline Saved all day articles, links, video 7:15 edition lead + sections Writer local + web tools Doctor 2nd model, on CPU The desks Finance desk 6am brief, pre-checked Knicks desk coverage, next games Music desk tonight's shows, 2 clips Manifest sections in order, one voice each On the 4070 Super 1. wait until 10 GB of video memory is free 2. the voice model reads it, 50 words at a time 3. music clips level matched and faded in 4. one loudness pass, out to mp3 Whisper transcribes the mp3 diff against the script, coverage per section, weakest chunk flagged mp3 posted to Discord, copy to iCloud with headlines, coverage numbers and a spot-check timestamp Any step that dies sends a message saying which one. A dead desk skips its segment, the show still ships.
Left to right is the writing, top to bottom is the rendering. Everything in blue happens on the card.

The voice trick

A 12 GB card is enough because of one trick. The voice model is Dia, 1.6 billion parameters, open weights, and it does five to twenty seconds of audio per call. So a ten minute show gets rendered in 50 word chunks and stitched, and the voices drift between chunks.

The fix: render one short anchor exchange with a fixed seed, once. Every chunk after that is generated as anchor plus new text, with the anchor audio as the prompt, then the anchor gets trimmed off the front. Each correspondent has its own anchor, so the Knicks guy always sounds like the Knicks guy.

The anchor trick Anchor rendered once, fixed seed Chunk text about 50 words Voice model anchor as the prompt anchor chunk audio trimmed off The show, stitched chunk 1 chunk 2 chunk 3 next Same anchor in front of every chunk, so the voice stays put. Each correspondent has an anchor of their own. Music clips are level matched to the speech, then one loudness pass over the whole thing.
Why a 1.6 billion parameter voice model can read ten minutes without wandering off.

The music desk

The music desk Venue emails their own mailbox Concert database artist, venue, date Rank tonight scored 0 to 100 Taste model rebuilt every Sunday from my records, my listening, my thumbs Playlist swap up to 10 songs, daily Two clips on air plus who's playing where thumbs a thumbs up or down on a pick changes next Sunday's model
The loop. The dashed line is what makes it get better over time.

If you wanted to build one

A few things I would tell you if you were starting one of these:

If you're running something like this at home, curious what you've got it doing. Shoot me a note.

Jim