Edge AI on the ESP32: What Actually Runs on a Microcontroller

Cover image for the edge AI and TinyML on ESP32 guide from eARgle Innovation Labs.

“AI on a microcontroller” sounds like marketing until you look at what it replaces: a device that had to stream audio or video to a server, over a network it cannot rely on, to answer a question as small as “was that a knock?”. Edge AI is not about intelligence. It is about not sending the data.

This is a practical map of what genuinely runs on an ESP32 in 2026 — which chip, which frameworks, what the memory ceiling really is, honest latency figures, and the three application classes that work well versus the ones that will waste your semester.

What “edge AI” means on a $5 chip

A microcontroller neural network is a fixed set of multiply-accumulate operations over a few tens of kilobytes of weights. There is no training on the device, no model bigger than the flash, no language model. What you get is a classifier: a function that turns a window of sensor data into one of a handful of labels, plus a confidence score.

That is a narrow capability, and it is transformative for exactly one reason — the comparison is not against a data centre, it is against the alternative embedded design:

Stream to a serverClassify on the device
Latency80–800 ms, network dependent5–150 ms, deterministic
Works offlineNoYes
PowerRadio on continuously — tens of mARadio off; wake only on an event
PrivacyRaw audio or video leaves the roomOnly the label leaves, if anything does
Running costBandwidth + cloud inference per deviceZero
Accuracy ceilingHigh — full-size modelsLimited — this is the real trade
The honest framing

You are trading accuracy for autonomy. A TinyML keyword spotter recognising eight commands will not match a phone assistant. It will run for months on a battery, answer in 30 ms, and work in a basement with no network — and for a light switch, a fall detector or a machine-fault alarm, that is the better device.

Which ESP32 can do it

ChipVerdict for on-device MLDetail
ESP32-S3The one to useDual LX7 at 240 MHz with vector (SIMD) instructions for the MAC-heavy inner loops, plus PSRAM support for image buffers. Espressif’s own ESP-DL and ESP-SR libraries target it, and Edge Impulse’s optimised kernels use its vector unit
ESP32 (classic)Workable for sensor models240 MHz dual-core, 520 KB RAM, no vector instructions. Accelerometer and audio-feature models run; images are painful
ESP32-C3Small models onlySingle RISC-V at 160 MHz, 400 KB RAM. Fine for accelerometer gesture or anomaly models of a few tens of kilobytes; not for audio spectrograms at speed
ESP32-C6 / C5Connectivity firstTheir advantage is Wi-Fi 6, Thread and Matter, not compute. Use them when the ML is trivial and the networking is the point
ESP32-P4Serious vision, with a caveatDual RISC-V at 400 MHz with hardware H.264 — the most capable of the family for video, but it has no radio and needs a companion chip, and Arduino support is still thin

If you are choosing hardware for a project with any ML in it, the answer is an ESP32-S3 with PSRAM. The price difference against a C3 is small; the difference in what fits is not. For the wider board decision, see our ESP32 vs ESP8266 vs Arduino Uno R4 comparison.

What actually works — three classes

1. Motion and vibration classification (the easiest win)

An accelerometer or IMU sampled at 50–100 Hz, windowed into one- or two-second frames, is a tiny input. Models land at 10–40 KB and infer in single-digit milliseconds even on a C3. This class includes gesture recognition, fall detection, gait analysis, and predictive maintenance from motor vibration — where a model learns the healthy spectrum and flags deviation.

This is where we point students first: the data is cheap to collect, the model is small, the accuracy is genuinely good, and you can verify it by shaking the board.

2. Audio event detection and keyword spotting

Not speech recognition — event and keyword detection. A digital MEMS microphone feeding an I²S peripheral, features computed as MFCC or mel-spectrogram, a small convolutional model over that. Realistic scope: 5–15 keywords, or a handful of sound classes such as glass breaking, a smoke alarm, a baby crying, a bearing screeching.

On an ESP32-S3 this runs comfortably in real time with a wake-word gate — a tiny always-on model that only lets the bigger one run when something sounds relevant, which is what keeps average power low.

3. Small-image vision (with limits stated up front)

96 × 96 or 160 × 160 grayscale, a handful of classes: person present / not present, a full versus empty bin, a gauge needle position, six product defect types. On an ESP32-S3 with PSRAM expect roughly 200 ms to 1 s per frame depending on model size — perfectly adequate for a sensor that decides once a second, useless for anything demanding a video frame rate.

What does not work

Face recognition across a population (as opposed to face detection); reading arbitrary text; a language model of any size; multi-object detection on a busy scene at frame rate; training or fine-tuning on the device. If a project brief needs one of these, the microcontroller is a sensor and a camera feed, and the model belongs on a Raspberry Pi class device or a server.

The memory and latency budget

The constraint that decides everything is RAM, and specifically tensor arena — the scratch memory TensorFlow Lite Micro needs for intermediate activations, which is separate from the model’s weights in flash.

Model typeWeights (flash)Arena (RAM)Latency, ESP32-S3 @240 MHz
Accelerometer gesture, dense net10–30 KB5–15 KB1–5 ms
Vibration anomaly (autoencoder)15–50 KB10–25 KB3–10 ms
Keyword spotting, small CNN30–90 KB30–60 KB20–60 ms per window
96×96 grayscale person detection250–350 KB70–140 KB200–500 ms
160×160 multi-class vision400 KB–1 MB150–350 KB (PSRAM advisable)0.5–2 s

Two rules follow from that table. First, quantise to int8 — it cuts model size roughly fourfold against float32 and is usually faster, at a typical accuracy cost of under a percentage point. Second, leave headroom: the Wi-Fi stack alone wants tens of kilobytes, so a model whose arena fills your remaining RAM will work beautifully until the moment it connects to a network.

The toolchain, end to end

  1. Collect data on the target hardwareNot from a public dataset recorded on studio microphones. The model must see the same microphone, the same mounting, the same noise floor it will face in production. This step decides your accuracy far more than the model architecture does.
  2. Label and split honestlyHold back a test set recorded in a different session — different day, different placement. Splitting one recording randomly gives you a flattering number that collapses in the field.
  3. Train smallEdge Impulse is the fastest route (browser-based, handles feature extraction, exports an Arduino library) and Espressif’s ESP-DL or plain TensorFlow/Keras give more control. Either way the target is thousands of parameters, not millions.
  4. Quantise to int8 and measure on deviceThe on-device numbers are the only ones that count. Report latency and RAM from the board, not from the training tool’s estimate.
  5. Deploy with a gate in frontRun inference only when something plausible happened — a sound above a threshold, motion detected, a timer tick. An always-on model is almost always a power bug.
  6. Log the disagreementsStore the inputs where confidence was marginal. Those few kilobytes are your next training set, and the difference between a demo and a product.

Four mistakes that sink these projects

  • Training on data from other hardware. A keyword model trained on clean dataset audio and run on a cheap MEMS mic in a fan-noisy room will look broken. Collect on the target.
  • Ignoring the negative class. A model trained on eight keywords will confidently classify a cough as one of the eight. You need an explicit “noise” or “unknown” class, and it usually needs more samples than any keyword.
  • Reporting the training accuracy. 99 % on a random split of one recording session means nothing. The number worth quoting comes from a separate session — and it will be lower, which is the point.
  • Forgetting the power budget. Continuous inference at 240 MHz burns tens of milliamps. If the device is battery-powered, the architecture is: sleep → cheap threshold wake → inference → sleep. Design that first, not last.

Where this is heading

Three shifts worth tracking. First, vector-capable microcontrollers are becoming the default rather than the premium option — the ESP32-S3’s SIMD unit and its equivalents on Arm’s newer M-profile cores make yesterday’s “too big” models routine. Second, Matter and Thread change what an edge classifier is for: a device that decides locally and publishes only a label fits a mesh far better than one streaming raw data, so the ESP32-C6 class and on-device ML are natural partners. Third, the tooling keeps collapsing the gap — collecting data from a board in a browser and getting back a quantised Arduino library is now a same-afternoon exercise, which moves the hard part firmly back to where it belongs: the data.

For anyone working in assistive technology, this matters more than the benchmarks suggest. A device that recognises a gesture, a fall or a spoken command without sending audio or video anywhere is not just cheaper to run — it is the only version of that device many users would accept in their home.

Frequently asked questions

Can an ESP32 run a large language model?

No. Even heavily quantised small language models need hundreds of megabytes of memory and sustained bandwidth that a microcontroller cannot provide; an ESP32 has a few hundred kilobytes of RAM. What fits on an ESP32 is a classifier of a few tens to a few hundred kilobytes. If a project needs language understanding, the microcontroller’s job is to capture and gate — and to decide locally when to spend network and cloud resources at all.

ESP32-S3 or ESP32-C3 for a TinyML project?

S3, unless the model is tiny and the budget is the binding constraint. The S3’s dual 240 MHz LX7 cores include vector instructions that the optimised kernels use for the multiply-accumulate inner loops, and it supports PSRAM for image buffers. A C3 at 160 MHz with no vector unit handles accelerometer and small anomaly models well, but audio spectrograms and any vision work will disappoint.

Do I need Edge Impulse, or can I use plain TensorFlow?

Both work. Edge Impulse is much faster to a working device — it handles data collection from the board, feature extraction (MFCC, spectral features), quantisation and packaging as an Arduino library, and its kernels are tuned for ESP32 hardware. Plain TensorFlow Lite Micro, or Espressif’s ESP-DL, gives you full control over architecture and memory layout, which matters once you are squeezing the last kilobytes. Start with Edge Impulse to learn the shape of the problem, then move down a level if the constraints demand it.

How much accuracy do I lose by quantising to int8?

Typically under one percentage point for the small models used on microcontrollers, when quantisation-aware training or a representative calibration dataset is used. In return the model shrinks roughly fourfold and often runs faster because integer MACs map better to the hardware. Always evaluate the quantised model on the device rather than assuming the float accuracy carries over.

Is on-device ML actually more private?

Substantially, if the design is honest about it. When classification happens on the microcontroller, raw audio, video or motion traces never leave the device — only a label, and only when there is something to report. That removes the most sensitive data from the network entirely. It is not automatic: a device that also uploads “debug samples” has given the property back. Decide what leaves the device, and make it the label.

Start with the hardware in your hands

Our motion and sensor kits are the fastest route into on-device ML — an IMU, a board and the data-collection workflow above.

MPU6050 motion pack Explore the courses
AR
Arun Roshan

Founder of eARgle Innovation Labs and eARgle Technologies, and a PhD scholar at VIT Vellore researching an integrated IoT–AI framework for assistive devices. On-device inference is a core constraint of that work: assistive hardware has to answer immediately, offline, and without shipping a user’s audio anywhere.

1 Comment.

Comments are closed.