Contents
Introduction Methodology
— Chapter 1. The Soundscape of the Above-Ground (Street) Environment — Chapter 2. The Acoustic Territory of Transition (Vestibule, Escalator) — Chapter 3. Sound as a Disciplinary Practice (Platform and Carriage) — Chapter 4. Reverse Transition and Auditory Recalibration
Conclusion Bibliography Field Notes
Introduction
Cities are rarely listened to. Rather, they are endured—as background noise, as an inevitable accompaniment to the visual experience. We look at streets, at faces, at architecture, but our ears are usually occupied: by headphones, by calls, by internal monologue. Yet, sound does not merely accompany space—it organizes it, marks its boundaries, dictates behavior, and shapes collective bodily practices.
This study focuses on the acoustics of transition—how the soundscape and the mode of listening change as a person moves between the above-ground and underground in a major city. A single route was selected as the empirical basis, encompassing a street segment, a descent into the subway, the platform, the train car, and the return to the surface. This route represents not just a geographic trajectory, but an acoustic territory (Brandon Labbell), where each segment possesses its own sound structure, norms of listening, and distribution of the private versus the public.
Secondly, we draw upon Brandon Labell’s spatial concept of «Acoustic Territories.» Labell, continuing the tradition of Henri Lefebvre and Michel de Certeau, illustrates how sound organizes social space through listening dispositions: private, public, and semi-private. The subway system serves as an ideal case study in this regard—it belongs to the state (public), yet each car mimics a state of temporary privacy.
Thirdly, the work explores the dichotomy between hearing and listening. Hearing is a physiological process—the perception of sound by the body. Listening, conversely, is a culturally mediated, aesthetic, and ethical action that involves evaluation, selection, and response. The transition from the street to the subway requires shifting from hearing (the automatic noise filter) to listening (the conscious tracking of announcements, evaluating a neighbor’s volume, and making an ethical decision—to comment or remain silent).
The research methodology integrates field sound recording with phenomenological description and structural analysis. Recordings were taken in the evening on a weekday, using a portable recorder positioned at human ear height. The visual component includes photographs of the recording locations.
The goal of this work is not merely to describe the sounds of the subway and the street, but to demonstrate how sound constructs social boundaries—where one can speak loudly and where one must whisper; to whom the voice in the loudspeaker belongs; and why silence in the car is perceived as the norm rather than as an absence of sound. Ultimately, this research is about how the city shapes our hearing—and how we can resist that conditioning by consciously listening.
Methodology
Study Type The research is a case study of a single urban route in St. Petersburg, incorporating both above-ground and underground segments. The selection of this case is driven by its typicality within a major Russian city and its concentration of acoustic contrasts, which allows for the application of a wide spectrum of theoretical concepts from the course.
Field Methods The empirical basis comprises: Field sound recording—conducted by the author in May 2026. Recordings were made at human ear height (approximately 1.6 m from the surface), simulating a natural auditory perspective.
Photodocumentation—capturing recording points, sound sources, and the spatial characteristics of the route. Photographs serve as a visual complement to the audio and enable the correlation of sound with its spatial context.
Participant listening—the author was situated within the recorded environment as a regular passenger, without inducing changes in the behavior of those around him.
- Analytical Procedure The analysis of each field segment (corresponding to Chapters 1–4) involves three stages: Descriptive—what is heard? (identification of sources, temporal structure). Categorical—how does this relate to theoretical concepts (figure/ground, acoustic territory, listening mode)? Interpretive—why is this significant for understanding the culture of listening in the city?
Block 1. Soundscape of the Above-Ground (Street)
The first chapter focuses on analyzing the soundscape of the above-ground space—the street adjacent to the subway entrance. According to R. Murray Schafer’s terminology, open urban space constitutes an acoustic environment with low reverberation and high entropy: sounds do not reflect off vertical surfaces, decay rapidly, and «escape» upward. This creates a specific acoustic architecture where each sound is typically heard only once and with clear directionality.
In such a landscape, three categories distinguished by Schaefer are easily discernible:
— Key sounds (figure)— those that the listener consciously attends to (traffic signals, footsteps, voices); — Background sounds (ground)— the continuous, often unconscious foundation (the hum of a highway, wind, rustling); — Sound signals (signals)— short, informational sounds that often carry a warning (a bell, a whistle, a lock click).
From the perspective of Brandon Labell’s acoustic territories, the street in the evening is a public space with weak sonic marking. This means that formally anyone can make any sound (within the law), but practically, unspoken norms exist: loud conversation or music is perceived as an intrusion, whereas quiet domestic sounds (footsteps, jingle of keys, the click of a gate) are acoustically neutral—that is, they do not require ethical judgment.
A distinct characteristic of this field fragment is the temporal mode of recording: late evening (10:30 PM). According to Scheffers' concept, the soundscape is not static; it changes depending on the time of day, the day of the week, and the season. The evening soundscape is sparse: there are fewer human voices, traffic operates at a less intense level, and the overall sound density decreases. This allows for the perception of micro-events that would be swallowed by the noise mass during the day—a kind of acoustic microhistory of a single intersection.
Furthermore, this fragment fundamentally includes bodily sound—the author’s footsteps in heels. This disrupts the illusion of the «invisible listener, ” which is classic in objectivist acoustics. Following the phenomenological tradition (and in accordance with the lecture on subjective position in field research), we examine the sound of one’s own body as a legitimate part of the soundscape, rather than as an „artifact“ to be removed. In this case, the heels are not simply a distraction but a marker of physical presence and simultaneously a gendered sound (in culture, heels are associated with femininity, office attire, and public self-presentation).
The field recording was taken at the entrance of a popular St. Petersburg metro station. The time selection (10:30 PM) was intended to avoid excessive acoustic density and capture a more sparse—and thus analytically transparent—soundscape. Unlike the daytime rush hour, when human voices and traffic create an intractable sonic mass, at 10:30 PM only a few passersby are present. This allows for the differentiation of individual acoustic events and the tracing of their temporal structure (sequence, duration, pauses).
A notable characteristic of the recording is the sound of the author’s heels. This sound is clearly audible during movement and fades during moments of stillness, marking the binary opposition of «motion / repose.» For the author, this additional sound is not an intrusion but rather possesses analytical value, as the heels represent the natural, everyday practice of wearing shoes. From the perspective of Scheffer’s sound ecology, this is an example of «kitsch» (a sound that does not convey important information but marks human presence), while from a phenomenological viewpoint, it constitutes an audio-corporeal self-portrait.
Below is a chronology of sound events with their terminological interpretation.
6th second—the gate swings open. This sound can be classified as a low-intensity auditory signal: it is brief but informative (someone is entering or exiting). The gate is an acoustic boundary between the private (yard, home) and the public (street). The sound of its opening marks a transition.
8th second—a woman brushes the gate with a handbag that has a metal buckle. A short, metallic click is an example of impulsive noise (a percussive, rapidly decaying sound). It carries no semantic information, but it creates an acoustic micro-event layer—those sounds that usually don’t capture attention but contribute to the texture of daily life.
Following this, there is the quiet rustling of the street and a distinct sound of footsteps. The rustling is a background sound (ground) in Scheffer’s classification. It is continuous, lacks a clear source, and serves as the acoustic «zero point» against which the volume of other events is measured. The author’s heel-clicks, in contrast, are a key sound (figure)—intermittent, with a sharp attack.
23rd second—the whistle of a braking bus. This is not an emergency signal, but the characteristic sound of a pneumatic braking system. From Scheffer’s classification standpoint, it is an auditory signal, but an unintentional one (the driver is not honking; the technical system is making the sound). Such «semi-signals» are interesting because they exist on the boundary between figure and ground.
At 33 seconds—a cyclist passes by, producing a whirring sound from the wheels. This whirring is an example of a periodic sound with frequency modulation (the wheel spokes intersect a magnetic field or simply vibrate). In the urban context, this sound is becoming rarer, yielding ground to electric vehicles and scooters. Its presence is an acoustic trace of hybrid mobility (muscular power + technology).
At 41 seconds—the bus whistle sounds again, but louder this time. This demonstrates how heterogeneous the soundscape is: the same type of event is perceived differently depending on distance and geometry.
At 45 seconds—a woman tosses cans into a trash receptacle near the subway entrance. The sound of the falling cans is sharp, metallic, with a long decay (the cans clatter against each other). This is an example of an event sound, one that disrupts the smooth background and draws attention. From a cultural perspective, it is also the sound of consumer waste—cans that were recently commodities have become refuse. Acoustics captures material culture faster than visual analysis.
At 46 seconds—a strong gust of wind. Wind is an atypical sound source in the city: it is neither human, nor technological, nor animal. Wind demonstrates the acoustic materiality of air—its capacity to carry sound and to produce it simultaneously (airflow around obstacles, turbulence). It serves as a reminder that the soundscape is not only a social construct but also a natural one.
Throughout the entire fragment, there is a constant background hum of the street—rustling, distant drone, occasional voices. The author’s heels are faintly, yet steadily, audible in the background at times. Their periodic appearance creates a rhythmic structure—a pulse upon which other, random sounds are layered.
Why is it important?
First, it demonstrates the temporal variability of the soundscape (Schafer). The evening soundscape is sparse, allowing listeners to perceive micro-events—the click of a buckle, the whir of a bicycle, the clinking of jars. This challenges the notion of the city as monotonous noise: the city sounds differently at different times, and these variations hold cultural significance (the evening is a time for slowing down, returning home, and transitioning from the public to the private).
Second, the fragment illustrates the corporeality of listening and the phenomenological stance of the researcher. The author’s footsteps in heels are not an artifact but an analytical key. They remind us that the listener is always within the sonic environment, in motion, and making sounds. As discussed in the lecture on subjective positioning in field research, ignoring one’s own body is equivalent to constructing a false picture of «impersonal,» «pure» perception. The heels also introduce a gender dimension: the sound of women’s steps in public space has historically been marked differently than men’s (an expectation of silence, softness, «appropriate» presence).
Third, the fragment establishes a baseline acoustic norm—the sound environment considered «normal» for the street in the citizens' consciousness. Comparing this to subsequent chapters (the subway, the train car) will reveal how sound changes when transitioning into an enclosed space: reverberation is added, directionality disappears, and the institutional voice of an announcer emerges. The street, in this sense, is the zero point of sound—the reference point.
Fourth, the fragment is vital for understanding the ethics of listening. The woman with jars, the cyclist, the bus driver—none of them consented to the recording. While such consent is not legally required in public space, it remains an ethical question. Field recording always balances between documentation and intrusion. Capturing the sounds of others' bodies (footsteps, coughing, the clinking of jars) is an act of acoustic appropriation, and the researcher must be accountable for this.
Fifth, the fragment shows how sound demarcates the boundaries between the private and the public (Labell). The gate is an acoustic marker of transition. The trash can is a place where private waste (jars from the home or from someone’s hand) becomes a public sound. The street at 10:30 PM is a semi-private zone: formally public, yet nearly empty, which makes every sonic event (someone’s footstep, someone’s cough) feel more intimate than it would during the day.
Block 2. Acoustic Transition Zone (Lobby, Escalator)
Chapter Two explores boundary acoustic zones—the spaces between the street and the subway car. According to Brandon Labell’s spatial concept, these zones constitute acoustic territories of regime shift: it is here that listening switches from diffuse, automatic reception (characteristic of the open street) to disciplined, ethically weighted listening (which will dominate the train car).
Physically, these zones are marked by changes in acoustic parameters: hard reflective surfaces (walls, ceiling, floor) appear, reverberation takes hold, the natural soundscape (wind, distant rumble) vanishes, and it is replaced by the mechanical, monotonous drone of ventilation, escalators, and turnstiles. From Scheiffer’s perspective, this represents a transition from an open soundscape of high entropy to a confined acoustic space where sound energy accumulates.
A distinctive feature of this field block is the temporal discontinuity. It was 10:30 PM on the street—evening, a rarefied environment. In the vestibule and on the platform, the time of day ceases to be an acoustic factor: the underground space is always uniformly lit, always possesses a constant background hum, always feels «timeless.» This creates an effect of acoustic disorientation—the listener loses their anchor to circadian rhythms, relying solely on artificial cues (announcements, train sounds).
Also crucial in this chapter is the category of semi-private listening (Labell). In the vestibule and on the escalator, people occupy public space yet behave as if they are alone: they do not converse, do not make eye contact, and avoid sonic interaction. The exception is when the train arrives: the flow of passengers creates a temporary acoustic densification, and the volume level and number of voices rise sharply.
It is also important to note the change in vertical acoustics: the ratio of direct to reflected sound shifts while descending the escalator. Where we primarily heard direct sounds from sources on the street, reflections begin to dominate in the escalator shaft—footsteps boom, and announcements sound «metallic.»
Fragment 1: Escalator / Entrance Zone (transition from street to foyer)
The first segment captures the earliest phase of transition—the moment when the listener is still physically outdoors but has just opened the subway doors. Acoustically, this is the sound boundary where a sharp shift in spectral composition occurs: the live, broadband sounds of the street (wind, rustling, distant voices) are replaced by the mechanical, monotonous, mid-frequency hum of ventilation and escalators.
It is expected that this segment will feature an effect referred to in sound studies as the acoustic shock of transition—a brief disorientation as the ear adjusts to the new reverberation characteristics. It is also important to track how the social density of sound changes: the street was deserted (10:30 PM), but a flow of people suddenly appears in the vestibule—a train has arrived, and passengers are ascending to the surface. This demonstrates that the subway’s soundscape is independent of the time of day above; beneath the ground, there is its own autonomous temporality.
The recording begins with the sound of the street—that distinct, rarefied evening soundscape analyzed in the first chapter. At the fourth second, a subway door opens. This moment is acoustically marked as a bifurcation point: the street sounds (more vibrant, broadband, with wind and rustling) are abruptly replaced by a mechanical hum. The contrast is so pronounced that it can be called an acoustic threshold—the physical door becomes a metaphor for the transition between different sound worlds.
Next, the recording is dominated by a constant, monotonous drone—this is the ventilation system, the escalators, perhaps the lighting transformers. From Scheffer’s perspective, this is pure background noise (ground), but unlike the street background (wind, rustling), this background is artificial and invariant—it does not change throughout the day, it is not dependent on the weather, and it lacks natural fluctuations. This creates a sense of acoustic stasis: time seems to halt underground.
Against this drone, faint conversations of people can be heard—without specific topics, only fragments of speech that cannot be pieced together into a coherent text. This is an example of peripheral speech: speech not intended for the listener, carrying no information for them, but marking the presence of other bodies.
At the 37th, 47th, and 50th seconds, mechanical collisions are heard—the sounds of turnstiles, metal railings, bags bumping against handrails. These are impulsive noises (like the buckle click in the first chapter), but louder and sharper due to the acoustics of the confined space—in the vestibule, any short sound receives an echo and seems louder than it would on the street.
At the 52nd second, a key event occurs: a whistle is heard, and then the conversations become louder, allowing us to distinguish individual words. The reason is the arrival of the train. Passengers flow up to the surface in a unified stream, and the quiet, desolate place suddenly becomes crowded.
This moment is crucial for understanding the acoustic architecture of the subway: the underground space is not isolated from train traffic. The arrival of the train creates a pulsation of social density—periods of silence are replaced by periods of acoustic concentration, when dozens of people pass through turnstiles simultaneously, talking, laughing, and coughing. This pulsation is not tied to the time of day above ground—at 10:30 PM, the flow can be as dense as at 5:00 PM, provided the train arrives.


This segment is significant because it captures the acoustic autonomy of the metro system. Outside, it was quiet and deserted (10:30 PM), but inside the vestibule, it suddenly becomes crowded and noisy. This visual contrast is even stronger: the area below is always uniformly bright, regardless of the time of day. It feels as if we are moving from night into day—from natural time into artificial, machine time. This supports Labell’s thesis that acoustic territories possess their own temporality, which is not synchronized with the outside world.
Fragment 2: Turnstile and Escalator (Descending)
