To create an immersive soundscape, record a continuous ambience bed in spatial audio, layer spot effects, dialogue, and room tone over it, then mix for depth using level, placement, and reverb. The Nomono Sound Capsule captures the ambience bed and four voice tracks in a single pass, and Studio Cloud handles cleanup and export.
An immersive soundscape is a layered audio environment that puts the listener inside a place instead of describing the place to them. Dialogue and music get the attention in most productions, but the ambience underneath is what convinces the ear. A bustling cafe, a quiet forest, a storm rolling in from a distance: each one tells the listener where they are before anyone says a word.
Spatial capture is what separates immersive from merely atmospheric. A mono or stereo ambience sits flat behind the voices. An ambisonic recording keeps the full sphere of the original space, so a gull passing overhead actually passes overhead when you render to spatial audio or binaural headphones. We have run the same harbor scene both ways in listening tests at the office, and the spatial version is the one people describe as "being there" rather than "hearing a recording of there".
Every soundscape I build breaks down into the same four layers. Knowing what each layer contributes tells you what to record on location, which is where most productions come up short. The table below maps each layer to what it does and how the Sound Capsule captures it.
| Layer | What it adds | Sound Capsule capture |
|---|---|---|
| Ambience bed | Continuous sense of place: weather, traffic, crowd, wildlife. The foundation everything else sits on. | The Space Recorder's 8-microphone ambisonic array records a 4-track ambisonic WAV of the full environment. |
| Dialogue | The story itself. Clarity here is non-negotiable. | Up to four Stellar wireless lav mics, each recording its own mono WAV, isolated from the bed. |
| Spot effects | Specific events with narrative weight: a door, a siren, a kettle, footsteps. | Recorded close on a Stellar lav held near the source, or picked out of the ambisonic bed by direction. |
| Room tone | The "silence" of the space. Used to patch edits so cuts stay invisible. | Two minutes of the array running with everyone quiet, captured before packing down. |
One pass, five files: four mono voice tracks plus the 4-track ambisonic bed, all time-aligned because they came off the same recorder. On shoots where I used to run a separate ambience rig next to the dialogue kit, sync drift between the two was a recurring editing cost. A single-recorder capture removes that failure point.
The workflow below is the one we use on our own field productions, from planning through export. Each step stands alone, so you can drop in wherever your project currently sits.
Write down every sound the scene needs before choosing gear. List the location, the time of day, the events that must be heard, and the emotional register you want. A dawn market and a midnight market are different recordings of the same place. Scouting with your ears, even for ten minutes on a phone memo, tells you when the location is at its best and what you will need to fight, usually traffic or HVAC.
Place the Space Recorder at the listener's position, roughly head height, away from reflective walls where you can manage it. Clip a Stellar lav on each speaking participant, up to four. The whole Sound Capsule kit weighs 2.9 lb with its charging case, so "bring the spatial rig" stopped being a logistics conversation for us. On windy exteriors, use wind protection on the array and accept that a gale will still win; I lost a clifftop recording in Lofoten to wind I was sure the foam would handle. Foam did not handle it.
Roll on the environment alone for at least two continuous minutes before any dialogue, longer if the location loops slowly, like surf or passing trains. A continuous bed lets you extend or shorten the scene in the edit without audible seams. Resist the urge to stop recording between takes. Storage is cheap and the one moment you cannot re-stage, a church bell, a flock lifting off, always happens while the recorder is off.
Run the dialogue takes with the array still rolling. The Stellar lavs give you clean, isolated voice tracks while the ambisonic bed keeps the scene alive underneath, which means you mix the balance later instead of committing to it on location. For spot effects with story weight, get close: hold a lav near the door, the kettle, the toolbox. Close-mic’ed effects layered over a spatial bed read as detail. Effects buried in the bed alone read as mush.
Before anyone moves, call quiet and record two minutes of the empty space. Room tone is the patch material for every edit you will make, and no plugin reconstructs it as well as the real thing. Forgetting this step cost me half a day on a factory documentary; the machines had a hum you cannot fake, and we had to go back. The two minutes now live on my slate checklist.
Upload the recordings to Studio Cloud for enhancement and cleanup, then assemble: bed first, at a level where it reads but never competes, dialogue on top, spot effects placed where the story needs them, room tone patching the joins. Automate the bed down 3 to 6 dB under speech and back up in the gaps, and the scene breathes. Static levels are the most common tell of a rushed mix.
Render the ambisonic bed to your delivery format. The Sound Capsule's ambisonic files feed that render directly, and the same source files fold down cleanly to binaural for headphone listeners, which is how spatial audio for podcasts reaches most audiences today. Check the final mix on at least two systems, good headphones and an ordinary speaker, because a soundscape that only works on studio monitors does not work.
Ambience earns its place in a production by doing narrative work, and I hold every layer to that standard. Four jobs come up constantly across podcasts, short films, and audio drama.
Location and time come first. A cyberpunk street and an Arctic research station need no narration when the sound design establishes them. Foreshadowing is the second job: a distant siren, a creak in the boards, thunder under an otherwise calm conversation all set the listener's nerves before the script confirms anything. Transitions are the third, where a shift in ambience carries the audience across space or time faster than any spoken signpost. Emotional tone is the fourth and the least appreciated. An echoing warehouse and a warm kitchen can host the same dialogue and produce two different scenes.
Restraint matters as much as capture quality here. On one of our early documentary projects we layered so much atmosphere into a fishing-village scene that the client asked why the interview felt "far away". We covered the full set of production lessons from that period in what we learned making immersive content, and the shortest version is: the bed serves the voice, never the reverse.
Most ruined soundscapes trace back to a small set of avoidable errors, and I have made every one of them at least once. Monitoring on earbuds instead of isolating headphones hides wind rumble and handling noise until you are home. Recording the bed after the dialogue wraps, when everyone is tired and half the crew has started chatting, produces a bed full of your own crew. Trusting that you can "grab ambience from the takes" leaves you with beds interrupted by every line of dialogue.
Distance from the action is the subtler one. Recordists park the array too close to the loudest element because it feels important, and the resulting bed has no depth to reveal. Put the array where the listener stands in the story, and let loud things be loud at a distance. For the dialogue side of remote work, our guide to recording a podcast outside the studio covers the voice-first version of the same discipline.
Create an immersive soundscape by recording a continuous spatial ambience bed, layering dialogue, spot effects, and room tone over it, then mixing for depth with level automation and placement. The Nomono Sound Capsule captures the ambisonic bed and four isolated voice tracks in one pass, so every layer stays in sync from capture through the mix.
A spatial soundscape needs a multi-capsule spatial microphone for the ambience bed, isolated mics for voices, wind protection, and isolating headphones for monitoring. The Nomono Sound Capsule packages this as one kit: a Space Recorder with an 8-mic ambisonic array plus four Stellar wireless lavs, 2.9 lb including the charging case.
An ambisonic recording captures sound as a full sphere around the microphone rather than a flat left-right image, stored by the Sound Capsule as a 4-track ambisonic WAV. The format matters because you can re-aim, rotate, and render the scene after the fact, to Dolby Atmos, binaural headphones, or plain stereo, without re-recording anything.
Record at least two continuous minutes of ambience per location, and more when the environment changes slowly, like surf, rail traffic, or crowd swells. Longer beds give the editor room to extend scenes without audible loops. Recording the bed before dialogue starts, while the location is undisturbed, produces cleaner material than grabbing it at wrap.
Podcasts are one of the strongest uses for immersive soundscapes, because audio carries the entire experience. Narrative shows, documentaries, and travel formats gain the most, and spatial episodes fold down to normal stereo for listeners without headphones that support it. Nomono publishes spatial audio for podcasts through the same Sound Capsule and Studio workflow described here.
No. Binaural rendering delivers the spatial effect on any ordinary stereo headphones, which is how most listeners hear immersive podcast audio today. Dedicated spatial playback systems, like Dolby Atmos speaker setups or head-tracked earbuds, add precision, but a well-mixed binaural fold-down carries the sense of place on the earbuds people already own.
Keep dialogue clear by recording voices on isolated lav mics rather than pulling them from the ambience, then automating the bed 3 to 6 dB down under speech. The Sound Capsule records each Stellar lav as its own mono WAV, separate from the ambisonic bed, so the voice-to-scene balance stays adjustable through the entire mix.
The Nomono Premium plan costs $1,188 upfront for the first year, which includes the Sound Capsule hardware and 12 months of Studio Cloud Premium, then $99 per month after that. Current terms and the lower-priced Professional tier are listed on the Studio Cloud pricing page.
Yes. Studio Cloud's Audio/Video Sync, available on all plans since March 2026, aligns Sound Capsule recordings with your camera footage automatically. Field crews record the spatial bed and lav tracks on the Capsule, shoot picture separately, and let Studio handle alignment, which removes the clapboard-and-waveform matching step from the edit.
Use purpose-made wind protection on the microphone array, monitor on isolating headphones so you hear rumble the moment it starts, and record backup takes in nearby sheltered positions. Foam alone fails in strong gusts. When a location is simply too windy, capture it anyway for reference and schedule a return; wind damage in a bed cannot be repaired convincingly in post.
Enjoy free express shipping to the listed locations within 2–5 business days.