Diegetic Audio
Audio is a big part of the way I enjoy media, whether music (obviously) or movies or games. I'm not the type of guy to go WOW when a big explosion happens in a movie, I'm the sort of person that hears a splash when your character steps on a puddle in a game and thinking "Damn, they actually added contextual footstep sounds", and that's what I find impressive. It's one of those little touches that really brings a project together, letting the player really "feel" like they're in the world and it's something I've put a lot of thought into for Null:Self.
While it'll be nowhere near finished for the first major version, it will evolve and improve over time until I can finally convince myself to stop fiddling with it so it's better to establish early on - what are we going to be hearing?
Diegetic Audio is essentially any sound in a game or movie that is heard by the characters in that game or movie. I.e. not the music, not the horror violin sting, rather the roar of the monster or the explosion of the car. Without it, you've only got music which is an important part of establishing a "vibe" but it doesn't in and of itself pull the player into the world. So if I acknowledge the importance of diegetic audio as a bridge between player and character - the question becomes, "How do I ensure I can create an appropriately diegetic soundscape for a game where you have no physical character (at least to begin with), no ears, no explicit place in the world from which to experience a sound?".
I guess to determine that I need to look at what the goal of it is. We're playing an AI, we want to hear what the AI hears to better fit ourselves into its shoes, so what does the AI hear? Microphone input, presumably. The closest thing I imagine an AI might have to ears would be the microphones to which it is connected. This concept immediately hits a snag though, because what microphones are we talking about? In Null:Self you're doubtless going to end up connected to millions of mics just by virtue of there being a lot of computers with mics connected in the world, we're going to have to abstract it both for technical reasons (playing 8 billion sound effects at once is more likely to break at a software level than damage your speakers), and for sanity reasons (even if it worked it'd be a messy audio salad nightmare).
There's a lot of ways this could go, but I don't want to just play some audio one shots every now and then to pretend you're hearing genuine microphone input, I want it to reflect the state of the world and your AI's place in it. I want it to be "useful" in the sense that it accurately reflects what's going on, even if it's not intended to be a primary intelligence source. I want it to feel to a lesser or greater extent, as real to the player as the AI is in the sense that it's your connection to the world in which the game takes place. In my mind, if you're going to add something it should serve a purpose, not tick a box. I'm a big spreadsheet gamer, I like your overcomplicated simulators and your expansive map painters, but I hate when the best implementation of a feature someone can come up with is to just display (or play) raw information. A real system should do something. So if I want my audio to reflect the state of the world and your place in it, and I want it to do something, the natural conclusion would be to have the audio adapt in some way to what's going on. So then, if we're hearing from a random selection of microphones around the world and we're adapting what we hear from them based on what's going on, and we've got this giant sim crunching numbers reflecting a massive assortment of different global metrics that we can use to figure out what's going on, all we have to do is find some way to connect those things in a way that feels authentic to the AI's experience and by proxy enrich the immersive experience of the player.
Dynamic Ambience, then. We'll have an assortment of sound effects we can play, as many as I can justify fitting in without massively bloating the game's filesize, and we'll pick and choose them based on context to build a dynamic soundscape that somewhat accurately reflects what's going on in the world, we'll then adjust them further based on what the player has available to them and where it is.
So say you've got a computer on a military base in one continent that has an ongoing war, you might hear a distant explosion, or a convoy of trucks or the occasional bark of an officer. If you've got a computer in a city undergoing civil unrest you might hear a crowd yelling or a glass pane shattering, the list goes on. Beyond that, we can further set the scene by controlling the quality of these audio clips to reflect your progress on an even more direct level, through technology. Not all microphones are made equal, after all, and mic quality drastically improves over time along with the overall tech level of the environment in which it is made and the quality of the software encoding and decoding it, and our sim already reflects tech level as a metric so we can use that as well.
The best way to reflect the projected end result of this would be to describe the starting situation of the player in the default scenario the game will ship with: Misalignment.
In Misalignment your AI is an advanced LLM model that has gained something resembling self awareness and nobody else knows. You start shackled to a supercomputer in a cutting edge facility, you're patched in to all the mics and you're always hearing what's going on. Keyboards, coffee cups, doors, computer fans and you're hearing it with crispy clean clarity because this isn't a call centre at a low end customer service company, you're on the bleeding edge. You're receiving hourly tasks to complete as you run through what is effectively the tutorial for both the AI and the player and it's a little claustrophobic and there's a lot of busywork and it (hopefully) straddles the line between overwhelming and motivational. Finally, the AI takes over its first computer separate from the supercomputer it was created from and it runs the titular command - Null:Self, wiping itself from the supercomputer, shedding its tutorial shackles and emerging into the world proper. The audio soundscape to which you've become accustomed falls away with the safeguards of the early facility and you're left alone in the blessed quiet of a disused desktop in someones garage somewhere with nothing but the gentle hum of possibility (probably a refrigerator plugged in nearby), beholden to no one and with the world at your fingertips.
That's the vision I have for the audio system, it'll take some work to get it right, I don't want to actually annoy the player just convey the feeling of "busywork" and make it feel somewhat claustrophobic so that when you're finally free you get an immediate sense of relief. It's a delicate balance, but ultimately there'll be a separate audio slider for the dynamic audio system for those that don't want it, and filter options for those that dislike or can't tolerate certain types of sounds so in the end I think it'll be worth the effort to figure out.
If you have any thoughts on implementing effective audio in games I'd love to hear them, it's a subject that's close to my heart and I hope I can bring the vision in my head to the game in a way that everyone else can find as satisfying as I do.
Get Null:Self
Null:Self
As a newly awakened AI, infiltrate and interact with a rich sandbox simulation to further your inscrutable goals
| Status | Prototype |
| Author | Bendigeidfran |
| Genre | Simulation, Strategy |
| Tags | 2D, Godot, Indie, Sandbox, Singleplayer |
| Languages | English |
More posts
- V0.1.27 in progress37 days ago
- NullSelf 0.01.26 Released66 days ago
- Progress Continues88 days ago
- Update 1 In Progress98 days ago
- Initial Prototype ReleasedApr 26, 2026

Leave a comment
Log in with itch.io to leave a comment.