PHYSICAL AI · ON-DEVICE INTELLIGENCE

Teaching the Pokédex to See
Without the Internet

A MobileNetV2 neural network, trained by transfer learning and quantized to run entirely on the Arduino UNO Q's CPU — plus an on-device voice loop (whisper.cpp + eSpeak). This is the Physical AI at the heart of Q-dex: it sees, hears and speaks, mostly without the internet.

MobileNetV2 95.8% quantized accuracy 2.59 MB model int8 · fully offline + on-device voice
4Pokémon classes
95.8%Val accuracy (quantized)
2.59 MBOn-device model
160²Input resolution
0Internet required

The Physical AI: "Who's That Pokémon?"

Q-dex has two ways to recognise a Pokémon. One asks the internet. The other thinks for itself.

The Reverse Image Version uploads a photo and queries Google Lens — powerful, covering all 1025 Pokémon, but dependent on a network connection. The PhyAI Challenge mode is the opposite: a neural network we trained ourselves, running directly on the device's processor, that recognises four Pokémon — Bulbasaur, Charizard, Pikachu and Squirtle — from the live camera feed with no internet whatsoever.

This is the literal meaning of Physical AI: intelligence embedded in the physical device, computing locally, in real time, from its own sensors. When you point the camera and the name appears with no network round-trip, that answer was computed on the Arduino itself.

At a glance

A MobileNetV2 classifier, trained by transfer learning on a free Colab GPU, fine-tuned, then quantized to an int8 .tflite model of 2.59 MB. It runs on the UNO Q via the LiteRT (ai-edge-litert) runtime, classifying camera frames roughly twice a second at 160×160 resolution.

Why On-Device, Not Cloud

The reverse-image mode already works. So why train our own model at all?

Because a challenge about Physical AI is asking a specific question: can the device itself be intelligent? Offloading every decision to a cloud API sidesteps that question. Running a model locally answers it. Four concrete reasons drove the decision:

1 · No connection dependency

The reverse-image path fails the moment Wi-Fi drops. The on-device model has no such failure mode — it works in a field, a basement, or on a stage with no network, which is exactly where a "handheld" device needs to work.

2 · Instant, no round-trip

Cloud recognition means capture → upload → API → response — often seconds. The local model returns an answer in a fraction of that, fast enough to classify the live camera feed continuously rather than on a single button press.

3 · Zero per-use cost

Every cloud query costs money and quota. The on-device model, once trained, runs unlimited times for free.

4 · It proves the concept

Training and deploying a real neural network onto a constrained embedded Linux board — and having it actually work from the camera — is the demonstration the challenge is looking for.

The Dataset

Four visually distinct Pokémon, ~40 images each, gathered and organised into class folders.

The four classes were chosen deliberately: Bulbasaur (green/blue, quadruped), Charizard (orange, winged), Pikachu (yellow, bipedal) and Squirtle (blue, shelled) are visually separable by colour and silhouette, which is exactly what makes a small-data classifier viable. The dataset was structured so folder names become the class labels:

pokemon_data.zip
  bulbasaur/   (~40 images)
  charizard/   (~40 images)
  pikachu/     (~40 images)
  squirtle/    (~40 images)
PropertyValue
Classes4 (bulbasaur, charizard, pikachu, squirtle)
Images per class~40
Total images~160
Train / validation split80% / 20% (seed 42, reproducible)
Training images per class~32 after the split
Input resolution160 × 160 × 3 (RGB)

The honest constraint

~32 training images per class is a very small dataset. A network trained from scratch on this would simply memorise the photos. Two techniques make it work anyway: transfer learning (reusing a model that already understands images) and data augmentation (manufacturing variety from what little we have). Both are described below.

Architecture — MobileNetV2 Transfer Learning

We don't teach the model to see from scratch. We borrow a model that already can.

MobileNetV2 is a convolutional network designed specifically for mobile and embedded devices — small, fast, efficient. Pre-trained on ImageNet (over a million photos across 1000 categories), it already knows how to detect edges, textures, shapes and colours. We freeze all of that learned knowledge and train only a small classifier head on top that maps its understanding onto our four Pokémon. This is why the whole thing works on ~160 images instead of the millions you'd otherwise need.

The model, layer by layer

Input (160×160×3)
   │
   ├─ Data augmentation      (train-time only: flip, rotate, zoom, brightness, contrast)
   ├─ MobileNetV2 preprocess (scale pixels to [-1, 1])
   ├─ MobileNetV2 base       (ImageNet weights — frozen in phase 1)
   ├─ GlobalAveragePooling2D (collapse feature maps to a vector)
   ├─ Dropout(0.3)           (regularisation — fights overfitting)
   └─ Dense(4, softmax)      (the only fully-trained head → 4 class probabilities)
Design choiceValueWhy
Base networkMobileNetV2 (ImageNet)Built for embedded; strong features for free
Input size160 × 160Light enough for the UNO Q's CPU
PoolingGlobalAveragePooling2DCompact, fewer parameters than Flatten
Dropout0.3Prevents memorising the small dataset
HeadDense(4) softmaxFour class probabilities

Training — Two Phases

Trained on a free Colab T4 GPU, in two deliberate stages.

Before any training, data augmentation is applied so the model never sees the exact same image twice. With so few photos, this is what prevents memorisation:

RandomFlip('horizontal')
RandomRotation(0.15)
RandomZoom(0.15)
RandomBrightness(0.2)
RandomContrast(0.2)
1
Phase 1 — Frozen base (15 epochs max)The entire MobileNetV2 base is frozen. Only the new Dense head trains, learning to map existing features onto the four Pokémon. Fast, and safe from overfitting. Adam optimiser, learning rate 0.001.
2
Phase 2 — Fine-tuning (10 epochs max)The last ~20 layers of MobileNetV2 are unfrozen and trained with a much smaller learning rate (0.00001), letting the network adjust its higher-level features for Pokémon-specific shapes and colours without destroying what it already knew.
3
Early stoppingBoth phases watch validation accuracy and stop early if it plateaus (patience 5), restoring the best weights seen. The model never trains past its peak.
HyperparameterPhase 1Phase 2
Trainable layersHead onlyHead + last ~20 base layers
Learning rate0.0010.00001
Max epochs1510
OptimiserAdam
LossCategorical cross-entropy
Batch size8 (small dataset → small batches)

Results

The number that matters: 95.8% validation accuracy on the exported, quantized model.

Accuracy was measured not on the original float model, but on the final int8 model that actually ships on the device — 23 of 24 held-out validation images classified correctly. This is the honest figure, because it's the model that runs on the Arduino, quantization losses included.

MetricResult
Validation accuracy (quantized, on-device model)95.8% (23 / 24)
Held-out validation set20% of the data, never seen in training
Live camera accuracy (real-world)~60% — lower, and honestly so (see below)

Confusion matrix (validation set)

Where the model agrees and disagrees with the truth. The green diagonal is correct predictions; near-perfect separation across all four classes.

bul
cha
pik
squ
bul
6
0
0
0
cha
0
6
0
0
pik
0
0
5
1
squ
0
0
0
6

Why we report ~60% for live use, not 95.8%

The 95.8% is on clean validation photos. In the real world — a phone screen or a toy under room lighting, at odd angles, with a busy background — live accuracy is closer to 60%. We document both numbers on purpose: the validation figure proves the model learned, the live figure is the honest field performance. Hiding the gap would be the wrong kind of documentation. This is also why the on-device mode is presented as a demonstration, while the 1025-Pokémon reverse-image mode remains the primary recognition path.

Quantization — Making It Fit

A GPU-trained model is too heavy for an embedded CPU. Quantization is what shrinks it.

The trained model uses 32-bit floating-point weights. To run efficiently on the UNO Q's CPU, it's converted to 8-bit integers (int8) — roughly a 4× reduction in size and a large speed-up, at the cost of a little accuracy. Crucially, the converter needs real images to calibrate the integer ranges correctly:

converter = tf.lite.TFLiteConverter.from_keras_model(model)
converter.optimizations = [tf.lite.Optimize.DEFAULT]
converter.representative_dataset = representative_dataset  # real training images
converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS_INT8]
converter.inference_input_type  = tf.uint8
converter.inference_output_type = tf.uint8
PropertyValue
FormatTensorFlow Lite (.tflite)
QuantizationFull int8, calibrated on real training images
Model size2.59 MB
Input tensor[1, 160, 160, 3] · uint8
Output tensor[1, 4] · uint8 (per-class confidence)

Because quantization can cost accuracy, the exported model is re-validated after conversion — the 95.8% figure above is that post-quantization number, not the float model's.

Deployment on the UNO Q

From a Colab download to a live camera classifier on the device.

Two files travel together: pokedex_phyai.tflite (the model) and labels.txt (the class order — turns a class number back into a name). On the device, inference runs through LiteRT (ai-edge-litert), not full TensorFlow, which would overflow the board's limited storage.

1
Grab a camera frameOpenCV captures a BGR frame from the USB camera.
2
PreprocessConvert BGR→RGB, resize to 160×160, cast to uint8 — exactly matching the training pipeline.
3
Invoke the modelThe LiteRT interpreter runs the frame through the int8 network on the CPU.
4
Read the resultargmax of the 4 outputs gives the class; the value becomes a confidence score. Below a 60% threshold the model stays silent rather than guess.
5
Show it liveThe recognised name appears on-screen with its confidence, updating roughly twice a second — continuous, not one-shot.

One hard-won deployment lesson

The obvious runtime, tflite-runtime, has no install candidate for this ARM board. The working path is ai-edge-litert — Google's current LiteRT package — imported as from ai_edge_litert.interpreter import Interpreter. The app tries three runtimes in order and reports exactly which failed and why, so a silent load failure can never masquerade as a model problem.

PHYSICAL AI · PART TWO

The Other Half: Voice AI

Seeing is only half of it. Q-dex also listens and talks — and, like the vision model, the listening and speaking both happen on the device itself.

Point the camera and the Pokédex recognises a Pokémon. But press the talk button and you can also have a conversation with Professor Oak. That voice assistant is a chain of three AI stages, and two of the three run entirely on the Arduino UNO Q — only the language reasoning in the middle uses the network.

1
You speak — whisper.cpp transcribes (on-device)Your spoken question is recorded from the USB microphone and turned into text by whisper.cpp, running locally on the board. No cloud speech API.
2
Claude thinks — the answer is written (cloud)The transcribed text is sent to Claude, an AI language model, which writes Professor Oak's reply. This is the one step in the loop that uses the internet.
3
eSpeak speaks — Oak talks back (on-device)Oak's reply is spoken aloud on the board by the eSpeak engine through the USB speaker, and the lens LED blinks in time with his voice.

whisper.cpp — offline speech-to-text

What it is: whisper.cpp is a fast, lightweight C++ port of OpenAI's Whisper speech-recognition model. It converts recorded audio into text — entirely on the device, with no internet.

Why we use it: to answer a spoken question, the device first has to turn your voice into text. Running that locally (rather than a cloud speech API) means no round trip, no per-use cost, and the assistant keeps working offline. We use the tiny.en English model — the best accuracy-for-speed trade-off on this hardware: small enough to run quickly on the UNO Q, accurate enough for short questions. Audio is captured with arecord at 16 kHz mono via the plughw:1,0 ALSA device (the raw hw device returns 44.1 kHz and breaks recognition).

eSpeak — offline text-to-speech

What it is: eSpeak is a compact, open-source speech synthesiser that turns text into spoken audio, running locally on the board.

Why we use it: once Oak's reply comes back as text, eSpeak reads it aloud through the PAM8403 amplifier and USB speaker — no cloud voice service, works offline. A physical mute button (GPIO 10) cuts the audio instantly, mid-sentence, so the conversation feels natural and interruptible. While Oak talks, the WS2812 lens LED blinks teal, driven by the ESP32.

Claude — the brain behind Oak

What it is: Claude is an AI language model (an LLM). It is the component that actually writes what Professor Oak says.

Why it's here — and why it's the one cloud step: generating natural, in-character, context-aware answers is beyond what fits on the board, so this single stage runs in the cloud. Claude is trainer-aware — it's given the active trainer's profile, so Oak can reference your real XP, badges and progress rather than generic replies. The honest framing: in Q-dex's voice loop, hearing (whisper.cpp) and speaking (eSpeak) are on-device; only the thinking uses the network.

On-device vs cloud — the honest picture

Q-dex uses five AI components. Three run on the device: MobileNetV2 (vision), whisper.cpp (speech-to-text) and eSpeak (text-to-speech). Two use the cloud: Claude (Oak's language reasoning) and Google Lens (the 1025-Pokémon reverse-image search). Every step that senses or acts — seeing, hearing, speaking — happens on the Arduino itself; only the heaviest language and whole-dex-image reasoning reaches out to the network. That balance is the point: as much intelligence in the physical device as the hardware can genuinely hold.

Summary

StageWhat we did
Dataset4 classes × ~40 images, 80/20 split, seed 42
ModelMobileNetV2 transfer learning + custom Dense(4) head
Training2-phase (frozen → fine-tune), augmentation, early stopping, Colab T4
ExportFull int8 quantization, calibrated on real images → 2.59 MB .tflite
Accuracy95.8% quantized validation (23/24); ~60% live, reported honestly
DeploymentLiteRT on the UNO Q CPU, live camera, ~2 fps, fully offline
Voice — hearingwhisper.cpp (tiny.en) speech-to-text, on-device, 16 kHz via plughw
Voice — speakingeSpeak text-to-speech, on-device, USB speaker, GPIO 10 mute
Voice — thinkingClaude (LLM) writes Professor Oak's trainer-aware replies — the one cloud step

The full training notebook — every cell, runnable end-to-end on a free Colab GPU — is included in the repository as pokedex_phyai_train.ipynb.