19/09/2026
You’re speaking from a distance and your phone is recording every word. But what is the machine actually capturing? Here’s the engineering.

Sound is air pressure waves. When you speak, your vocal cords vibrate and push surrounding air, creating compressions and rarefactions that travel toward your phone. The microphone doesn’t capture sound directly. It captures these air pressure changes.

Modern smartphones use a MEMS microphone, Micro Electro Mechanical System, sometimes just around 1 mm × 1 mm. Inside is a microscopic silicon diaphragm, like a tiny drum, with a fixed backplate forming a capacitor. When sound hits it, the diaphragm flexes slightly, changing the capacitance and producing a voltage change. Your voice becomes a continuous analog signal.

But computers understand numbers, not continuous waves. That’s where an ADC, Analog to Digital Converter, comes in. It takes rapid snapshots of the analog signal, with each snapshot called a sample, and converts the voltage at that moment into a number.

This is where the Nyquist Theorem matters. The sample rate needs to be at least twice the highest frequency you want to capture. Human hearing goes up to around 20,000 Hz, so 44,100 samples per second became a standard sampling rate for digital audio.
Your phone is essentially taking thousands of these measurements every second and turning them into numbers.

The same ADC concept that converts joystick movement in a PS5 controller can convert your voice too.

44,100 measurements per second. Each one captures a moment of the sound wave. Together, they become your digital recording.