StoryCraft · Omni-Audio Framework
Omni-Audio Framework
Anatomy of every audio asset in the Omni system — podcasts, audiograms, voiceover, sonic branding. Audio is trust. Poor audio kills a great script; great audio elevates a mediocre one.
Universal Audio Principles
| Rule | Standard | Why |
|---|---|---|
| Source quality | Right mic, close, treated room | You cannot fix room echo in post. Capture clean. |
| Loudness target | Podcast −16 LUFS · Streaming −14 · Broadcast −23 | Platforms normalize; over-loud files get pulled down and sound worse. |
| Dynamic range | Compress speech to ≤6dB range | Listener shouldn't reach for the volume knob. |
| Silence | Cut breaths > 0.4s; keep pauses that carry meaning | Pacing is content, not filler. |
| Consistency | Same intro/outro, same voice tone, same mix across episodes | Trained ear = returning listener. |
Long-Form Podcast
Weekly show, interview, or narrative. Purpose: deep authority, community, and inbound trust. 20–60 minute runtime.
Episode structure
Every episode follows the same skeleton so listeners know what to expect.
Structure
- Cold open (10–30s) — best soundbite from the episode, pulled forward.
- Intro music + brand ID (5–10s) — show name, host tagline.
- Host welcome + episode setup (60–90s) — what listener will get.
- Segment 1 / interview body (main content).
- Mid-roll ad or CTA (30–60s) — optional.
- Segment 2 / closing insight.
- Outro (15–30s) — CTA to subscribe/review, tease next episode.
- Sonic sign-off — same music sting every time.
Gear
- Entry — Shure MV7, Rode PodMic USB. Solo or 2-host.
- Pro — Shure SM7B or Electro-Voice RE20 into Rodecaster or Focusrite.
- Headphones — closed-back (Sony MDR-7506, Audio-Technica M50x).
- Room — soft-treated (rugs, curtains, foam). Avoid parallel bare walls.
Mic Technique
- Mouth 4–6 inches from mic capsule.
- Slightly off-axis (10–15°) to reduce plosives.
- Pop filter mandatory on dynamic mics used close.
- Stay consistent — don't lean in/out mid-sentence.
Mix / Master Recipe
- Sample rate — 48 kHz / 24-bit record; export 44.1 / 16.
- Noise reduction — subtle (2–4 dB) — over-processed sounds hollow.
- EQ — HPF 80Hz, gentle cut 200–400Hz (mud), boost 3–5kHz (presence).
- De-esser — 5–8kHz, threshold to taste.
- Compression — 3:1 ratio, 4–6dB gain reduction.
- Limiter — ceiling −1dBTP.
- Loudness — normalize to −16 LUFS integrated (podcast standard).
- Export — MP3 128–192 kbps mono OR 192–256 kbps stereo.
Common Mistakes
Recording in a bare room
Echo destroys presence. Add soft surfaces before you spend on a mic.
Mixed loudness episode-to-episode
Listener reaches for volume. Master to −16 LUFS every time.
Long cold open explaining show format
Bury the hook. Start with the best moment of the episode.
Every guest sounds different
Levels + EQ + compression per guest, then match to host.
Short-Form Audio Clips & Audiograms
30–90 second pull quotes from long-form, published as video-with-waveform. Purpose: social distribution of podcast content.
Pull quote · Waveform · Captions
Structure
- Hook first (0–3s) — the strongest sentence, not the setup.
- Body (10–60s) — one complete thought.
- Payoff / punchline — the "aha" line.
- End tag (2s) — show name + where to find full episode.
Visual Spec
- Aspect — 1:1 (feed) or 9:16 (Reels / TikTok / Shorts).
- Waveform animation — bars or line, synced to audio.
- Captions — burned in, animated, 24pt+ bold.
- Show art — cover art top or corner.
- CTA overlay — "Full episode → [platform]".
Audio Prep
- Pull cleanest 30–60s from the episode.
- Trim breath in/out, remove ums.
- Re-normalize to −14 LUFS (streaming target).
- Add subtle music bed if voice-only feels sparse.
Distribution
- Post native to Reels, TikTok, Shorts, LinkedIn.
- Caption includes ep number + link.
- 3–5 clips per long-form episode = a week of social.
Voiceover — Video, Ads, Explainers
Purpose: guide, sell, teach. The voice IS the brand for the duration of the play.
Script pacing · Tone · Delivery
Script Rules
- Read aloud before recording — anything you stumble on gets rewritten.
- Short sentences. Contractions. Second person.
- Pace target — 150–170 words/minute for ads, 130–150 for tutorials.
- Mark stress + pauses with slashes: "This / changes / everything."
Recording
- Same mic setup as podcast; add pop filter always.
- Multiple takes per line — pick best syllable-by-syllable in edit.
- Match tone to visual: high-energy ad = up, testimonial VO = warm/measured.
- Room tone: record 30s of silence for edit fills.
Mix Recipe
- Voice sits −6dB against music at −24dB.
- Auto-duck music when VO plays.
- EQ: cut 250Hz mud, boost 4kHz presence.
- Compress 4:1, gain reduction 4–6dB.
- Loudness: match to video platform target (−14 for streaming ads, −23 for broadcast).
Delivery Types
- Announcer — high energy, front-of-mic, punchy.
- Conversational — closer, warmer, "friend talking."
- Narrator — measured, authoritative, documentary.
- Character — stylized, for entertainment/brand persona.
Sonic Branding — Jingles, Stings, Idents
Purpose: instant brand recognition through sound. Think Netflix "tudum" — 2 seconds, entire brand.
Memorable · Short · Ownable
Sting / Ident (1–3s)
- 3–5 notes maximum.
- Distinctive interval (perfect 5th, minor 3rd — pick one and own it).
- Same instrument palette across all uses.
- Ends on tonic / resolved — earworm effect.
Jingle (5–15s)
- Simple hook, lyrics reference brand name or promise.
- Melody within an octave range (singable).
- Repeatable across ad lengths (5s / 15s / 30s cut-downs).
Usage Rules
- Same sting at intro AND outro — bookend recognition.
- Never remix/redo per project — consistency is the whole point.
- Match visual brand palette (upbeat brand ≠ moody sting).
- Test it in isolation — recognizable without visual?
Where It Lives
- Intro of every podcast episode.
- Ident on every video (start + end card).
- Hold music, IVR menus.
- Event walk-on / walk-off cue.
Loudness & Format Cheat Sheet
| Format | LUFS Target | True Peak | Sample Rate / Bitrate | Notes |
|---|---|---|---|---|
| Podcast (mono) | −16 LUFS | −1 dBTP | 44.1 kHz / 128 kbps MP3 | Apple + Spotify standard |
| Podcast (stereo) | −16 LUFS | −1 dBTP | 44.1 kHz / 192–256 kbps | Only if music-heavy |
| Streaming ad / audiogram | −14 LUFS | −1 dBTP | 48 kHz / 192 kbps AAC | Spotify/YouTube target |
| Broadcast (TV/radio) | −23 LUFS | −1 dBTP | 48 kHz / WAV | EBU R128 / CALM Act |
| Video voiceover | Match host video −14 | −1 dBTP | 48 kHz / WAV embed | Duck music under VO |
| Sonic sting | −14 LUFS | −1 dBTP | 48 kHz / WAV | Deliver stems for placement |