Parent info
Parts you need
Affiliate links — we may earn a small commission
Try this circuit in your browser!
Run the code, press the buttons and watch what happens — before you buy any parts. No account needed.
Open in Simulator →Your music makes the lights dance. Every beat, every note, in real time.
Imagine this: you stick a 60-LED strip behind your monitor and plug in a small mic. You hit play on any music. The strip explodes — deep red pulses from both ends on every kick drum. Green blooms outward from the center on vocals. A blue shimmer sweeps the whole strip on every hi-hat.
That’s FFT — Fast Fourier Transform — made visible. You’re watching one of the most important algorithms in computing run 50 times per second and turn math into light.
All for $15.

What you’ll need
| Part | What it does | Price |
|---|---|---|
| ESP32-S3-DevKitC-1 | Brain. Runs FFT and controls LEDs simultaneously. | ~$8 |
| WS2812B LED strip, 1m / 60 LEDs | The visual output. Each LED individually addressable. | ~$8 |
| INMP441 MEMS microphone | Digital mic — captures audio as a number stream over I2S. Powered at 3.3V only. | ~$3 |
| 5V 2A power supply | 60 LEDs at full brightness need more current than a USB port provides. | ~$5 |
Total: ~$15 | Time: ~1.5 hours | Difficulty: ●●○○○
You’ll also need: A 330Ω resistor for the LED data line. A 100µF capacitor across 5V/GND near the LED strip (prevents power spike damage on startup). Short lengths of wire.
How it works (60 seconds)
Think of it like a tiny scientist listening to music with superhuman ears.
The INMP441 microphone captures audio as a digital number stream — 44,100 numbers per second, each representing air pressure at that moment. The ESP32 collects 256 of those numbers, then runs FFT math on them.
FFT takes 256 mixed-up numbers (the raw sound wave) and produces 128 output numbers where each output number answers one question: “how loud is the sound at this specific frequency right now?” Output 1 answers “how loud is the bass at ~172 Hz?” Output 64 answers “how loud are the treble frequencies at ~11,000 Hz?”
You group those outputs into three bands (bass, mids, highs) and map them to LED colors and positions. Bass = red from both ends. Mids = green from center. Highs = blue everywhere. Update 50 times per second. Done.
Step 0: Handle the microphone first
Time: ~5 minutes
The INMP441 is a tiny chip on a breakout board with small pads. Solder your wires before anything else — it’s easier without other components in the way.
INMP441 pins to identify:
- VDD — power (3.3V — absolutely not 5V, it will burn out)
- GND — ground
- L/R — channel select (tie to GND for left channel)
- SCK — I2S bit clock
- WS — I2S word select (frame sync)
- SD — I2S data output (sound data to ESP32)
Use thin wire (22–26 AWG) and a fine-tipped soldering iron. Verify each connection with a multimeter’s continuity mode before proceeding.
Step 1: Wire it up
Time: ~15 minutes
INMP441 Microphone (5 wires — 3.3V power):
- INMP441 SCK → GPIO 4
- INMP441 WS → GPIO 5
- INMP441 SD → GPIO 6
- INMP441 VDD → 3.3V — red wire, absolutely not 5V
- INMP441 GND → GND — black wire
- INMP441 L/R → GND (selects left channel)
WS2812B LED Strip (3 wires): 7. LED DIN → 330Ω resistor → GPIO 13 — orange wire 8. LED VCC → 5V power supply (not from USB/board 5V pin) — red wire 9. LED GND → GND (shared with board) — black wire
Power protection: 10. 100µF capacitor across 5V and GND at the LED strip’s power input (any polarity: the + leg to 5V)
Check: INMP441 is on 3.3V. LED strip is on dedicated 5V. 330Ω resistor on the data line. 100µF cap on LED power. These protection components prevent the most common hardware damage.
Common mistake: Powering the INMP441 from 5V. The chip maxes at 3.6V — 5V kills it instantly. Always 3.3V.
Step 2: Install libraries and flash
Time: ~10 minutes
Install in Arduino IDE Library Manager:
- FastLED — WS2812B LED control
- arduinoFFT — FFT analysis (search “arduinoFFT”)
Select your ESP32 board and upload the code:
The big picture first. Every loop of this program does three things:
- Listen: the microphone captures 256 audio samples — a tiny slice of sound.
- Analyze: FFT math splits that slice into 128 frequency answers — “how loud is the bass right now? the mids? the highs?”
- Draw: those three loudness numbers map to red, green, and blue patterns across 60 LEDs.
Repeat 50 times per second and you get a light show that dances with the music.
// ========== CHOOSE YOUR BOARD ==========
// Uncomment the line for YOUR board:
#define BOARD_S3 // ESP32-S3-DevKitC-1
//#define BOARD_C6 // ESP32-C6-DevKitC-1
// ========================================
#ifdef BOARD_S3
#define PIN_NEOPIXEL 13
#define PIN_MIC_SCK 4
#define PIN_MIC_WS 5
#define PIN_MIC_SD 6
#endif
#ifdef BOARD_C6
#define PIN_NEOPIXEL 5
#define PIN_MIC_SCK 0
#define PIN_MIC_WS 19
#define PIN_MIC_SD 20
#endif
#include <FastLED.h>
#include <driver/i2s.h>
#include <arduinoFFT.h>
#define LED_PIN PIN_NEOPIXEL
#define NUM_LEDS 60
#define MIC_SCK PIN_MIC_SCK
#define MIC_WS PIN_MIC_WS
#define MIC_SD PIN_MIC_SD
#define FFT_SAMPLES 256
#define SAMPLE_RATE 44100
CRGB leds[NUM_LEDS];
double vReal[FFT_SAMPLES], vImag[FFT_SAMPLES];
ArduinoFFT<double> FFT(vReal, vImag, FFT_SAMPLES, SAMPLE_RATE);
void setupMic() {
i2s_config_t cfg = {
.mode = (i2s_mode_t)(I2S_MODE_MASTER | I2S_MODE_RX),
.sample_rate = SAMPLE_RATE,
.bits_per_sample = I2S_BITS_PER_SAMPLE_32BIT,
.channel_format = I2S_CHANNEL_FMT_ONLY_LEFT,
.communication_format = I2S_COMM_FORMAT_STAND_I2S,
.intr_alloc_flags = ESP_INTR_FLAG_LEVEL1,
.dma_buf_count = 4,
.dma_buf_len = FFT_SAMPLES,
.use_apll = false,
};
i2s_pin_config_t pins = {
.bck_io_num = MIC_SCK,
.ws_io_num = MIC_WS,
.data_out_num = I2S_PIN_NO_CHANGE,
.data_in_num = MIC_SD
};
i2s_driver_install(I2S_NUM_0, &cfg, 0, NULL);
i2s_set_pin(I2S_NUM_0, &pins);
}
void readMic() {
int32_t rawBuf[FFT_SAMPLES];
size_t bytesRead = 0;
i2s_read(I2S_NUM_0, rawBuf, sizeof(rawBuf), &bytesRead, portMAX_DELAY);
for (int i = 0; i < FFT_SAMPLES; i++) {
vReal[i] = (double)(rawBuf[i] >> 8);
vImag[i] = 0.0;
}
}
float getBandLevel(int startBin, int endBin, float scale) {
float sum = 0;
for (int i = startBin; i <= endBin; i++) {
sum += vReal[i];
}
float avg = sum / (endBin - startBin + 1);
return constrain(avg / scale, 0.0f, 1.0f);
}
void setup() {
FastLED.addLeds<WS2812B, LED_PIN, GRB>(leds, NUM_LEDS);
FastLED.setBrightness(100);
setupMic();
}
void loop() {
readMic();
FFT.windowing(FFTWindow::Hamming, FFTDirection::Forward);
FFT.compute(FFTDirection::Forward);
FFT.complexToMagnitude();
float bass = getBandLevel(1, 3, 800000.0f);
float mids = getBandLevel(4, 23, 400000.0f);
float highs = getBandLevel(24, 64, 200000.0f);
bass = pow(bass, 0.6f);
mids = pow(mids, 0.6f);
highs = pow(highs, 0.6f);
int bassCount = (int)(bass * 15);
int midsCount = (int)(mids * 20);
int highBright = (int)(highs * 60);
for (int i = 0; i < NUM_LEDS; i++) {
CRGB color = CRGB::Black;
if (i < bassCount || i >= (NUM_LEDS - bassCount)) {
color += CRGB(200, 10, 0);
}
int center = NUM_LEDS / 2;
if (abs(i - center) < (midsCount / 2)) {
color += CRGB(0, 180, 20);
}
color.b = max(color.b, (uint8_t)highBright);
leds[i] = color;
}
FastLED.show();
}
Line-by-line: what every line does and why
Lines 1–3: Borrowing three instruction books
#include <FastLED.h>
#include <driver/i2s.h>
#include <arduinoFFT.h>
FastLED knows how to talk to WS2812B LEDs. driver/i2s.h enables the I2S audio hardware inside the ESP32. arduinoFFT.h provides the Fast Fourier Transform math. All three are installed once in Arduino IDE’s Library Manager.
Lines 5–11: Constants
#define NUM_LEDS 60
#define FFT_SAMPLES 256
#define SAMPLE_RATE 44100
NUM_LEDS = 60 — one metre of LED strip at 60 LEDs per metre. FFT_SAMPLES = 256 — how many audio measurements we collect before running the math. More samples = finer frequency detail, but more time. SAMPLE_RATE = 44100 — the mic captures 44,100 numbers per second, each representing air pressure at that instant. Same quality as a CD.
Lines 13–16: Creating the LED array and FFT buffers
CRGB leds[NUM_LEDS];
double vReal[FFT_SAMPLES], vImag[FFT_SAMPLES];
ArduinoFFT<double> FFT(vReal, vImag, FFT_SAMPLES, SAMPLE_RATE);
leds[60] is a shelf with 60 compartments — each holds one colour (R, G, B values). vReal and vImag are the two arrays FFT needs. Think of vReal as “the actual sound samples” and vImag as “a helper array the math uses internally — starts at zero.” ArduinoFFT creates the analysis engine, connecting it to both arrays.
setupMic(): Configuring the digital microphone
The INMP441 microphone sends audio over I2S — a three-wire digital audio protocol. cfg.sample_rate = 44100 tells it to capture 44,100 samples per second. bits_per_sample = 32BIT — the mic sends 32-bit numbers even though only the top 18 bits contain useful audio (the bottom bits are noise). dma_buf_count = 4 creates four memory buffers that fill up in rotation — while the CPU processes one buffer, the hardware fills the next. This prevents gaps in recording.
readMic(): Filling the FFT buffer
i2s_read(I2S_NUM_0, rawBuf, sizeof(rawBuf), &bytesRead, portMAX_DELAY);
for (int i = 0; i < FFT_SAMPLES; i++) {
vReal[i] = (double)(rawBuf[i] >> 8);
vImag[i] = 0.0;
}
i2s_read waits until 256 audio samples are ready (blocking — like waiting at a tap until the glass fills). Then the loop copies them to vReal. >> 8 is a right shift by 8 bits — it discards the bottom 8 noisy bits and shifts the useful 24 bits down. vImag[i] = 0.0 resets the imaginary part to zero before each FFT run — required by the math.
getBandLevel(): Measuring one frequency range
float avg = sum / (endBin - startBin + 1);
return constrain(avg / scale, 0.0f, 1.0f);
After FFT runs, vReal contains 128 numbers — each answers “how loud is sound at this frequency right now?” Bins 1–3 cover ~172–516 Hz (bass). We average the bins in the range and divide by scale to normalise to 0.0–1.0. constrain(x, 0, 1) clips the result like a guardrail — never below 0 or above 1.
loop() — The FFT pipeline
FFT.windowing(FFTWindow::Hamming, FFTDirection::Forward);
FFT.compute(FFTDirection::Forward);
FFT.complexToMagnitude();
Three steps, always in this order. windowing applies a Hamming window — it gently fades the audio samples at the edges of our 256-sample chunk to zero. Without this, the sudden start and end of each chunk create fake high-frequency artefacts. compute runs the actual Fourier transform. complexToMagnitude converts the raw FFT output into amplitude values we can use.
loop() — Gamma correction
bass = pow(bass, 0.6f);
pow(x, 0.6) is a power function — like a gentle curve. Without it, only the loudest sounds ever push the LED bars up. With it, quiet sounds still look vivid. Your eyes work the same way — they’re more sensitive to small changes in dark areas than bright ones. This is called gamma correction and your TV uses it too.
loop() — Drawing the LEDs
if (i < bassCount || i >= (NUM_LEDS - bassCount)) {
color += CRGB(200, 10, 0);
}
int center = NUM_LEDS / 2;
if (abs(i - center) < (midsCount / 2)) {
color += CRGB(0, 180, 20);
}
color.b = max(color.b, (uint8_t)highBright);
Bass glows from both ends toward the middle — i < bassCount catches the left end, i >= NUM_LEDS - bassCount catches the right end. Mids bloom from the centre outward — abs(i - center) is the distance from the centre LED; we light those within midsCount/2 of it. Highs add blue brightness to every LED using max — “use whichever blue value is higher, the existing one or the highs value.” += on a colour mixes the two colours together. Finally FastLED.show() pushes all 60 colour values to the strip in one fast burst.
The whole thing in one sentence
Every loop: record 256 audio samples, run FFT to separate bass/mids/highs, use those three loudness levels to paint a red/green/blue pattern across the strip, and push it to the LEDs 50 times per second.
First thing to try: Upload and play music near the microphone. If nothing lights up, clap loudly right next to the mic. Then adjust the scale values in getBandLevel — halving the bass scale makes bass LEDs react to quieter sounds.
Check: After upload, the LEDs should immediately react to sound near the mic. Tap the table, clap, play music. If nothing happens, check mic wiring first.
Step 3: Tune the sensitivity
Time: ~5 minutes
The three scale values control how much signal makes the LEDs light up fully:
- Bass:
800000.0f - Mids:
400000.0f - Highs:
200000.0f
If bass LEDs never light up with music: halve the bass scale value (try 400000.0f). If bass LEDs are always at full brightness: double the scale (try 1600000.0f). Tune each band until the visualization feels alive but not maxed out.
Step 4: Mount it!
Time: ~10 minutes
Behind your monitor: Peel the adhesive backing on the LED strip. Stick along the back edge of your monitor facing the wall. The reflection creates indirect lighting that changes with your music.
Frosted acrylic frame: Cut a 60mm-wide strip of 3mm frosted acrylic to the same length as your LED strip. Mount the strip 15mm behind the acrylic. Individual LED dots become a smooth glowing gradient through the diffuser.
Aluminum channel: LED strip suppliers sell aluminum extrusion channels with frosted diffuser caps. Slide the strip in, push the cap on. Professional housing, zero effort.
Play music near the mic. Watch the strip react. Every kick drum pulses red from both ends. Every vocal line blooms green from the center. Every hi-hat shimmer adds blue across the whole strip.
What just happened (what you learned)
-
FFT (Fast Fourier Transform) — invented in 1965, the most used algorithm in computing. Your phone uses it to compress audio (MP3). WiFi routers use it to decode wireless signals. MRI machines use it to reconstruct brain images. Here it turns 256 audio samples into a frequency spectrum — separating “loud at 200 Hz” from “loud at 4000 Hz” in a single calculation.
-
DMA (Direct Memory Access) — why the I2S setup uses
dma_buf_count = 4. Normally, the CPU copies every byte from hardware to RAM manually. DMA lets hardware write directly to RAM while the CPU does something else — here, running FFT and updating LEDs. Without DMA, the ESP32 would spend all its time shuttling audio bytes and have nothing left for visualization. -
Bin resolution —
sample_rate / FFT_samples = 44100 / 256 = 172 Hz per bin. More samples = finer frequency resolution but more calculation time. 256 samples takes about 1.5ms to compute. 1024 samples would be finer but 4× slower. -
Non-linear scaling —
pow(x, 0.6)is gamma correction — the same math your monitor uses to display images correctly. Audio magnitudes span enormous ranges: a whisper might be 100, a kick drum might be 10,000,000. Without gamma correction, only the loudest sounds ever look bright.
Level Up
Change the color scheme. Replace red/green/blue with fire palette (bass = orange-yellow), cyan (mids), white pulse (highs). Or use CHSV(hue, 255, brightness) and map frequency energy directly to hue.
Add peak-hold. Real spectrum analyzers hold the peak position for half a second before falling. Track peakBass, peakMids, peakHighs. If current level exceeds peak, update it; otherwise decay slowly (peak *= 0.99). Draw peak as a single bright LED above the bar.
Add a silence animation. When all band levels drop below 0.05 for 2+ seconds, trigger a slow color wave screensaver. Snap back to visualizer mode when music returns.
★★ You completed: Audio Visualizer LED Strip!
Troubleshooting
| Problem | Fix |
|---|---|
| LEDs don’t react to sound | Check INMP441 wiring: SCK→4, WS→5, SD→6, VDD→3.3V. Check L/R pin is to GND. Try clapping loudly right next to the mic. |
| LEDs react but bass is always full | Increase bass scale value (try 2000000.0f). You may be in a noisy room. |
| LEDs flicker randomly in silence | This is normal — FFT picks up room noise. Add a minimum threshold: only light LEDs if band level > 0.05f. |
| First LED always on | Missing 330Ω resistor on data line. Add it between GPIO 13 and LED DIN. |
| LEDs very dim | Check LED VCC is 5V from a real supply, not 3.3V from the board. |
| ESP32 resets or crashes | 60 LEDs at full white draw 3.6A. Use a dedicated 5V 2A supply for the LEDs, not the USB port. |