Lijo Jacob
May 10, 2021

A new lightweight CNN model for Automatic Speech Command Recognition on Microcontrollers

Abstract

The need for Automatic Speech Command Recognition (ASCR / ASR) on IoT devices is gaining traction because of the increased interest in non-touch-based applications. This article introduces a new lightweight convolutional neural network (CNN) for ASR on microcontrollers. The proposed model is comparable to current state-of-the-art networks with a very low parameter count of less than 63k. The new model gives an accuracy of 96.13% on the Google Speech Commands V2 dataset. A comparative study of results on previous models on the same dataset is also presented.

1. Introduction

Currently, ASR is being done using human-computer interfaces like Google speech, Siri, Alexa which require a client-server mode for its operation as the neural network is very computationally expensive. So, for devices with no internet access, doing speech recognition becomes a non-viable option, as the device itself cannot run very computationally expensive networks. In this paper, we present a very lightweight neural network suitable for low-power devices like microcontrollers.

The network should satisfy the following constraints to run on microcontrollers:

Memory Footprint: Very small memory footprint, ranging in the few 10s of KiloBytes
Low Compute Power: Limited MCPS processor cores in the range of a few 10s of MHz to sub-200MHz.
Offline: All processing should be done locally without cloud connectivity

The remaining contents of this paper are organized in the following sections: 2: Model Architecture, 3: Training Methodology, and 4: Results.

2. Model Architecture

The model is composed using Keras with a Tensorflow backend. An audio file consists of a single word; hence, the current model can be thought of as a classification model. Each audio file is a single channel mono wav file, sampled to 16000 Hz and provided as the input to the network. 40 band Mel Frequency Cepstral Coefficients (MFCC) features are extracted from the audio sample and fed into the network. These MFCC features are fed into a custom convolutional neural network (CNN) for generating the classification results.

Fig 1.0 High-Level Model Architecture

3. Experimental Results

3.1 Model Integration

For all the experiments, the github repository [3] was referred. To maintain uniformity of all experiments, all aspects of the repository except for our custom model was kept the same.

3.2 Experimental Setup

For the experiment, the dataset used is the Google Speech Commands (GSC) 12 class set with the following keywords: ‘_unknown_’, ‘left’, ‘on’, ‘stop’, ‘right’, ‘off’, ‘down’, ‘up’, ‘no’, ‘go’, ‘yes’, ‘_silence_’ [1]. All keywords are sampled at 16kHz and are of duration 1s.

The GSC V2 comprises 36 folders with the dataset split into train, validation, and test based on predefined percentages. 10% of the total dataset is split as a test and 10% as validation, the remaining 80% is categorized as train data. The keywords not belonging to the above-mentioned keyword list are classified as unknowns. The composition of the train and test set is as shown in the table below.

Class	Counts	Class	Counts
on	3086	on	396
right	3019	right	396
stop	3111	stop	411
up	2948	up	425
down	3134	down	406
no	3130	no	405
go	3106	go	402
left	3037	left	412
yes	3228	yes	419
off	2970	off	402
unknown	6154	unknown	816
silence	3077	silence	408

Table 1.0: a) train data counts per class b) test data counts per class

The background noise class present in the Google Speech Commands dataset is not considered for training as a class, but it is mixed with other speech signals to create augmented data. Silence class is generated by multiplying a random file with zeros and the count of silence class is calculated as 10% of total files in any random folder. All metrics and methods are in accordance with standard practices [1,2]. The table below is generated with reference from [1].

Model	Accuracy	Model Size (Kbits)
DNN	90.6	3576
CNN+strd	95.6	4232
CNN	96.0	4848
GRU(S)	96.3	4744
CRNN(S)	96.5	3736
SVDF	96.9	2832
DSCNN	96.9	3920
TinySpeech-A	94.3	127
TinySpeech-B	91.3	53
LMU1	96.9	1683
LMU2	95.9	361
LMU3	95.0	105
LMU4	92.7	49
IGN-CNN(our model)	96.13	490

Table 2.0 Accuracy Results for different networks, based on GSC V2 dataset

Fig 2.0 Scatterplot comparison of various networks: Accuracy vs Model size

4. Conclusion

A new, lightweight CNN-based model for ASR, optimized for embedded microcontroller devices, was developed. We have benchmarked the model against comparable models using the Google Speech Commands V2 dataset. The accuracy results and total model footprint are comparable to the prevalent state-of-the-art models. This model architecture has been deployed on multiple variants of low-cost microcontrollers from leading semiconductor manufacturers

5. References

Peter Blouw, Gurshaant Malik, Benjamin Morcos, Aaron R. Voelker, and Chris Eliasmith “Hardware Aware Training for Efficient Keyword Spotting on General Purpose and Specialized Hardware”, https://arxiv.org/pdf/2009.04465.pdf
Oleg Rybakov, Natasha Kononenko, Niranjan Subrahmanya, Mirko Visontai, Stella Laurenzo, “Streaming keyword spotting on mobile devices”, https://arxiv.org/pdf/2005.06720.pdf
Reference code, https://github.com/google-research/google-research/tree/master/kws_streaming

39 thoughts on “A new lightweight CNN model for Automatic Speech Command Recognition on Microcontrollers”

Sign up to get 100 USDT
June 17, 2025 at 8:47 pm

I don’t think the title of your article matches the content lol. Just kidding, mainly because I had some doubts after reading the article.
Edgardex
July 15, 2025 at 4:12 pm

pharmacie en ligne suisse sans ordonnance: pharmacie viagra generique – pharmacie en ligne cialis
Michaeleveme
July 16, 2025 at 4:35 pm

https://zorgpakket.shop/# online apotheek goedkoper
ScottAverm
July 16, 2025 at 5:16 pm

24/7 apotek [url=https://tryggmed.shop/#]glyserol apotek[/url] munnbind apotek
Altonneody
July 16, 2025 at 5:31 pm

stiv sГҐle apotek: kjГёlebind apotek – lusekur apotek
KennethTeeks
July 16, 2025 at 6:22 pm

europese apotheek: online apotheek gratis verzending – medicatie online bestellen
Michaeleveme
July 16, 2025 at 10:07 pm

https://snabbapoteket.shop/# apotek bestÃ¤lla medicin
Williamtreve
July 16, 2025 at 11:39 pm

https://tryggmed.com/# sГёndags apotek
ScottAverm
July 16, 2025 at 11:42 pm

a-vitamin syre krem apotek [url=https://tryggmed.com/#]apotek ГёyedrГҐper[/url] apotek retinol
Altonneody
July 16, 2025 at 11:43 pm

betrouwbare online apotheek zonder recept: holland apotheke – online apotheek goedkoper
Michaeleveme
July 17, 2025 at 3:37 am

https://snabbapoteket.com/# pipetter apotek
Altonneody
July 17, 2025 at 5:47 am

skalpell apotek: Snabb Apoteket – mina recep
ScottAverm
July 17, 2025 at 6:02 am

snorking apotek [url=https://tryggmed.shop/#]kondomer apotek[/url] peppermynteolje apotek
Michaeleveme
July 17, 2025 at 9:40 am

https://snabbapoteket.com/# apotek djurrecept
Williamtreve
July 17, 2025 at 11:41 am

http://zorgpakket.com/# aptoheek
Altonneody
July 17, 2025 at 12:58 pm

tabletter: SnabbApoteket – stetoskop apotek
ScottAverm
July 17, 2025 at 1:37 pm

serum apotek [url=http://snabbapoteket.com/#]nГ¤tapotek sverige[/url] mens tabletter
Michaeleveme
July 17, 2025 at 4:30 pm

https://zorgpakket.com/# medicijnen online bestellen
Altonneody
July 17, 2025 at 8:17 pm

nГ¤t apotek: SnabbApoteket – antikroppstest apotek
KennethTeeks
July 17, 2025 at 8:20 pm

munnbind med ventil apotek: TryggMed – tattoo krem apotek
ScottAverm
July 17, 2025 at 9:23 pm

arkaden apotek [url=https://tryggmed.com/#]TryggMed[/url] dagkrem apotek
Michaeleveme
July 17, 2025 at 11:19 pm

https://zorgpakket.shop/# afbeelding medicijnen
Williamtreve
July 18, 2025 at 3:28 am

https://snabbapoteket.shop/# hГ¤stfoder online
Altonneody
July 18, 2025 at 3:54 am

mobil apotek: SnabbApoteket – apotek menskopp
KennethTeeks
July 18, 2025 at 4:20 am

ibuprofen flytande barn: glucosamin apotek – gravid app frukt
ScottAverm
July 18, 2025 at 5:35 am

nattlysolje apotek [url=http://tryggmed.com/#]svarte munnbind apotek[/url] collagen plus apotek
Michaeleveme
July 18, 2025 at 6:42 am

http://zorgpakket.com/# online doktersrecept
Altonneody
July 18, 2025 at 11:27 am

apteka holandia: apteka amsterdam – medicijnen bestellen apotheek
KennethTeeks
July 18, 2025 at 11:58 am

apotheek medicijnen bestellen: farma – online apotheek gratis verzending
Michaeleveme
July 18, 2025 at 1:26 pm

https://snabbapoteket.shop/# apotek pÃ¥ internet
Davidphara
July 18, 2025 at 2:37 pm

MediMexicoRx [url=http://medimexicorx.com/#]order kamagra from mexican pharmacy[/url] rybelsus from mexican pharmacy
sign up for binance
July 18, 2025 at 5:58 pm

Thank you for your sharing. I am worried that I lack creative ideas. It is your article that makes me full of hope. Thank you. But, I have a question, can you help me?
Vernonagind
July 18, 2025 at 6:03 pm

https://expresscarerx.org/# ExpressCareRx
Bobbynew
July 18, 2025 at 6:40 pm

viagra pills from mexico: mexican pharmacy for americans – MediMexicoRx
Robertfluot
July 18, 2025 at 7:31 pm

https://medimexicorx.com/# MediMexicoRx
Davidphara
July 18, 2025 at 8:44 pm

online pharmacy china [url=https://expresscarerx.online/#]ExpressCareRx[/url] domperidone uk pharmacy
LewisBex
July 18, 2025 at 9:11 pm

pharmacy uk: ExpressCareRx – propranolol uk pharmacy
Robertfluot
July 19, 2025 at 12:48 am

http://expresscarerx.org/# ExpressCareRx
Davidphara
July 19, 2025 at 2:41 am

safe place to buy semaglutide online mexico [url=http://medimexicorx.com/#]buy propecia mexico[/url] order from mexican pharmacy online

Cookie	Duration	Description
cookielawinfo-checkbox-analytics	11 months	This cookie is set by GDPR Cookie Consent plugin. The cookie is used to store the user consent for the cookies in the category "Analytics".
cookielawinfo-checkbox-functional	11 months	The cookie is set by GDPR cookie consent to record the user consent for the cookies in the category "Functional".
cookielawinfo-checkbox-necessary	11 months	This cookie is set by GDPR Cookie Consent plugin. The cookies is used to store the user consent for the cookies in the category "Necessary".
cookielawinfo-checkbox-others	11 months	This cookie is set by GDPR Cookie Consent plugin. The cookie is used to store the user consent for the cookies in the category "Other.
cookielawinfo-checkbox-performance	11 months	This cookie is set by GDPR Cookie Consent plugin. The cookie is used to store the user consent for the cookies in the category "Performance".
viewed_cookie_policy	11 months	The cookie is set by the GDPR Cookie Consent plugin and is used to store whether or not user has consented to the use of cookies. It does not store any personal data.

info@ignitarium.com

A new lightweight CNN model for Automatic Speech Command Recognition on Microcontrollers

Abstract

1. Introduction

2. Model Architecture

3. Experimental Results

3.1 Model Integration

3.2 Experimental Setup

Table 1.0: a) train data counts per class b) test data counts per class

Table 2.0 Accuracy Results for different networks, based on GSC V2 dataset

Fig 2.0 Scatterplot comparison of various networks: Accuracy vs Model size

4. Conclusion

5. References

39 thoughts on “A new lightweight CNN model for Automatic Speech Command Recognition on Microcontrollers”

Leave a Comment

Stay informed

NEWS & VIEWS

Join our team

APPLY

PRIVACY POLICY

©2025 Ignitarium Technology Solutions, All Rights Reserved

Newsletter

An ISO 9001:2015 certified company

Great Place to Work® Certified

We are a leading provider of Product Engineering Services, offering expertise in Semiconductor design, Multimedia & Imaging, Connectivity, Cloud & Enterprise solutions, and Machine Learning & Deep Neural Networks

Semiconductor

Software

Ecosystem

Resources

Contact Us

Request for Video

info@ignitarium.com

A new lightweight CNN model for Automatic Speech Command Recognition on Microcontrollers

Abstract

1. Introduction

2. Model Architecture

3. Experimental Results

3.1 Model Integration

3.2 Experimental Setup

Table 1.0: a) train data counts per class b) test data counts per class

Table 2.0 Accuracy Results for different networks, based on GSC V2 dataset

Fig 2.0 Scatterplot comparison of various networks: Accuracy vs Model size

4. Conclusion

5. References

39 thoughts on “A new lightweight CNN model for Automatic Speech Command Recognition on Microcontrollers”

Leave a Comment

Stay informed

NEWS & VIEWS

Join our team

APPLY

PRIVACY POLICY

©2025 Ignitarium Technology Solutions, All Rights Reserved

Newsletter

An ISO 9001:2015 certified company

Great Place to Work® Certified

We are a leading provider of Product Engineering Services, offering expertise in Semiconductor design, Multimedia & Imaging, Connectivity, Cloud & Enterprise solutions, and Machine Learning & Deep Neural Networks

Semiconductor

Software

Ecosystem

Resources

Contact Us

Human Pose Detection & Classification

Features:

Target Markets:

OCR / Pattern Recognition

Use cases :

Highlights :

Behavior Monitoring

Use cases :

Highlights :

Attire & PPE Detection

Use cases :

Use cases :

Request for Video

Real Time Color Detection​

Use cases :

Highlights :

Missing Artifact Detection

Use cases :

Highlights :

Real Time Manufacturing Line Inspection

Use cases :

Highlights :

Ground Based Infrastructure analytics

Use cases :

Highlights :

Aerial Analytics

Use cases :

Highlights :

SANJAY JAYAKUMAR

Request Free Demo

RAMESH EMANI

​Manoj Thandassery

MALAVIKA GARIMELLA​

PRADEEP KUMAR LAKSHMANAN

SONA MATHEW

ASHWIN RAMACHANDRAN

AZIF SALY

RAJU KUNNATH

PRADEEP SUKUMARAN

SUJEET SREENIVASAN

RAJIN RAVIMONY

SIBY ABRAHAM

SUDIP NANDY

SUJEETH JOSEPH

SUJITH MATHEW IYPE

RAMESH SHANMUGHAM

Real Time Color Detection

Manoj Thandassery

MALAVIKA GARIMELLA