1. First principle: sound controls intelligibility, space, attention, time, and emotion
SOUNDTRACK
= dialogue
+ narration
+ production sound
+ ambience
+ room tone
+ effects
+ Foley
+ music
+ silence
+ spatialization
+ dynamics
+ loudness
+ transitions
Audio is not subordinate cleanup. It can carry continuity across picture changes, create off-screen space, reveal causality, and determine whether a podcast remains understandable.
2. Sound-object vocabulary
DIALOGUE_ISO
MIX_TRACK
NARRATION
ADR
WILD_LINE
ROOM_TONE
AMBIENCE
HARD_EFFECT
FOLEY
DESIGNED_EFFECT
RISER
IMPACT
DRONE
MUSIC_STEM
MUSIC_FULL_MIX
STINGER
IDENT
SILENCE_EVENT
REFERENCE_TONE
3. Audio-event dimensions
source/entity
start/end time
semantic role
story function
foreground/background
on_screen/off_screen
diegetic/non_diegetic
perspective
position/azimuth/elevation/distance
width
frequency balance
dynamic envelope
transient character
reverb/space
noise state
phase/polarity
loudness contribution
rights/provenance
4. Dialogue signal chain
A configurable dialogue chain:
INGEST
→ CHANNEL_SELECTION
→ SYNC
→ PHASE/POLARITY CHECK
→ CLIP REPAIR
→ DENOISE/DEHUM/DECLICK AS NEEDED
→ EDIT
→ ROOM-TONE FILL
→ GAIN STAGING
→ SPECTRAL SHAPING
→ DYNAMICS
→ DE-ESS
→ MATCH SPEAKERS/TAKES
→ BUS PROCESSING
→ LOUDNESS CONTROL
→ TRUE-PEAK CONTROL
→ QC
Order may change by material. Processing should solve a diagnosed problem, not be applied merely because a preset exists.
5. Dialogue-editing dimensions
intelligibility
naturalness
noise floor
room consistency
proximity
breath policy
mouth noise
plosives
sibilance
level consistency
tonal consistency
cross-talk
latency/drift
edit audibility
Rules:
IF noise reduction creates warbling or speech loss
THEN reduce processing, use spectral repair locally, fill with room tone, or select another source.
IF two microphones capture the same speaker with delay
THEN select one source or align carefully; do not leave uncontrolled comb filtering.
IF a dialogue cut changes room tone abruptly
THEN bridge with matched room tone or ambience.
IF a breath carries emotion or phrasing
THEN preserve it; otherwise reduce only as needed.
6. Perspective grammar
CLOSE/INTIMATE
NATURAL_CONVERSATIONAL
DISTANT
OFF_SCREEN_NEAR
OFF_SCREEN_FAR
THROUGH_DEVICE
THROUGH_WALL
PUBLIC_ADDRESS
MEMORY/SUBJECTIVE
Perspective is built from:
direct-to-reverberant ratio
high-frequency loss
pre-delay
level
stereo position
early reflections
noise/environment
occlusion filtering
Sound perspective should agree with picture unless deliberate counterpoint is declared.
7. Ambience and room-tone grammar
AMBIENCE
= place identity
+ time identity
+ activity
+ weather
+ density
+ perspective
+ continuity bed
+ notable events
Ambience functions:
orient place
bridge cuts
indicate time
signal danger or safety
create scale
mask edits
foreshadow change
show absence through change
Room tone is captured or synthesized per setup and microphone perspective. Do not use one generic noise bed across visibly different spaces.
8. Sound-effect grammar
Effects can be:
LITERAL — heard source shown or implied
ENHANCED_LITERAL — recognizable source with controlled emphasis
SYMBOLIC — sound represents a concept/emotion
TRANSITIONAL — bridges or punctuates structural change
INTERFACE — confirms system action
HYPERREAL — heightened detail beyond ordinary perception
SUBJECTIVE — heard through a character’s state
For every effect, identify:
cause
onset
material
size
perspective
space
impact
resonance
decay
consequence
9. Music grammar
Music dimensions:
narrative function
diegetic status
motif
key/mode
harmony
tempo
meter
rhythmic density
instrumentation
timbre
register
dynamic contour
phrase structure
edit points
stem availability
rights
Music functions:
IDENTITY
ORIENTATION
MOMENTUM
TENSION
RELEASE
EMPATHY
IRONY
COUNTERPOINT
TRANSITION
MEMORY/CALLBACK
SCALE
CLOSURE
Rules:
IF music tells the audience what to feel before the evidence earns it
THEN delay, reduce, or remove music.
IF dialogue carries vulnerable testimony
THEN protect intelligibility and consider silence or restrained texture.
IF a music cue changes function
THEN align the change with a story turn, visual turn, or explicit counterpoint.
IF edits need flexibility
THEN use stems, loops, alternate endings, and phrase-aware cuts.
10. Motif grammar
MOTIF
= identifiable sound/music unit
+ associated entity/idea/state
+ recurrence
+ controlled variation
Variation dimensions:
tempo
instrumentation
register
harmony
density
reversal
fragmentation
spatial treatment
A callback becomes meaningful when the motif returns under changed story conditions.
11. Silence grammar
Silence may function as:
ATTENTION_RESET
EMOTIONAL_SPACE
ABSENCE
SHOCK
ANTICIPATION
BOUNDARY
REALISM
CONTRAST
RESPECT
“Silence” can contain controlled room tone. Sudden digital zero may sound like a fault unless that rupture is intended.
12. Audio-transition grammar
J_CUT
L_CUT
PRELAP
POSTLAP
AMBIENCE_BRIDGE
MUSIC_BRIDGE
SOUND_MATCH
SOUND_MOTIF
RISER_AND_IMPACT
REVERB_TAIL
FILTER_TRANSITION
PERSPECTIVE_TRANSITION
HARD_SILENCE
CROSSFADE
13. Podcast voice matching
Match across speakers and sources by evaluating:
integrated level
short-term level
tonal center
low-frequency proximity
sibilance
room signature
noise floor
stereo/mono width
reverb
compression character
Do not force every voice into identical timbre. The goal is coherent listening effort, not erased individuality.
14. Loudness grammar
Loudness targets are delivery profiles, not universal creative mix targets.
Store:
integrated_loudness_target
loudness_range_target
maximum_true_peak
measurement_standard
channel_layout
gating behavior
normalization policy
Reference profiles included in this package:
EBU_BROADCAST_REFERENCE: -23 LUFS integrated, profile-specific true-peak limit
ONLINE_PODCAST_STARTING_POINT: -16 LUFS integrated, configurable by platform/workflow
The final profile must follow the distributor or broadcaster’s current specification. Keep a dynamic master before platform normalization.
15. Spatial audio grammar
channel_layout
object position
azimuth
elevation
distance
spread
divergence
room model
head tracking status
binaural render
fold-down behavior
For stereo podcasts:
keep essential dialogue center-compatible
use width for ambience/music without compromising mono
check phase and fold-down
16. Audio repair grammar
DECLIP
DECLICK
DECRACKLE
DEHUM
DENOISE
SPECTRAL_REPAIR
PLOSIVE_REPAIR
DEESS
DEREVERB
PHASE_ALIGN
TIME_STRETCH
PITCH_REPAIR
DRIFT_CORRECTION
DROPOUT_PATCH
ROOM_TONE_FILL
ADR_OR_PICKUP
Repair hierarchy:
source replacement
→ local edit
→ targeted restoration
→ alternate take
→ pickup/ADR
→ transparent disclosure if necessary
17. Mixing automation dimensions
clip gain
fader automation
pan
EQ automation
compression/ducking
reverb send
music ducking
ambience perspective
mute state
transition curves
Automation should follow semantic and spatial events, not only waveform level.
18. Caption and transcript relation
Audio events relevant to understanding should have textual representation:
speaker identity
meaningful non-speech sound
music when narratively relevant
off-screen source
language changes
inaudible/uncertain segment
Caption wording and timing are part of audio QA, not merely delivery decoration.
19. LAKA sound variations
Baseline
clean, level, match, and preserve natural space
Minor
change music density, ambience level, or local transition
Major
new motif, subjective perspective, or sound-led montage treatment
Structural
rebuild sequence around audio prelaps, silence, or parallel sonic spaces
Paradigm
picture-led scene becomes audio-first experience or interactive spatial soundscape
20. Audio failure modes
DIALOGUE_UNINTELLIGIBLE
OVER_RESTORATION
ROOM_TONE_HOLE
COMB_FILTERING
CLIPPING
TRUE_PEAK_OVER
LOUDNESS_PROFILE_MISMATCH
MUSIC_MASKS_SPEECH
MUSIC_MANIPULATES_EVIDENCE
AMBIENCE_LOCATION_MISMATCH
PERSPECTIVE_MISMATCH
PHASE_OR_MONO_FAILURE
REMOTE_DRIFT
UNLICENSED_MUSIC
CAPTION_MISREPRESENTS_SOUND
21. Audio acceptance test
Can every essential word be understood at ordinary listening level?
Do edits preserve natural phrasing and room continuity?
Does sound perspective agree with space or declared subjectivity?
Does music perform an earned narrative function?
Are silence and ambience intentional?
Does the mix meet the selected loudness/true-peak profile?
Does mono/fold-down remain intelligible?
Are rights, stems, captions, and provenance complete?