All problem statements
SIH26172HardwareSmart Automation

Low Latency and Efficient Voice Activator for Edge Devices

Indian Space Research Organisation(ISRO)

Ideas submitted
88 / 500
Deadline
30 September 2026
Category
Hardware
Theme
Smart Automation

Looks like it needs

IoT / EmbeddedChatbots / Voice

Background As voice-controlled IoT proliferate, processing everything in the cloud is too costly, privacy-invasive, and slow. The future belongs to hybrid architectures where the edge handles the initial 'wake-up' and the cloud handles the heavy lifting.

Description Build an ultra-lightweight, highly accurate keyword spotting (KWS) model that runs locally on a low-power device. Upon detecting the keyword, the system must instantly and efficiently stream the subsequent audio to a remote Automated Speech Recognition (ASR) server with minimal data overhead and latency.

Key Metrics for Evaluation

• Efficiency: Model size (RAM/Flash footprint) and CPU usage during idle listening. • Accuracy: High true-positive rate for the keyword with near-zero false activations. • Latency: The time delta between the keyword ending and the cloud ASR receiving the audio stream.

Software & Framework Restrictions

• Open-Source Only: The use of proprietary, closed-source, or commercial voice-activation SDKs is strictly prohibited. • Allowed Frameworks: Teams must build their keyword spotting (KWS) pipelines using open-source machine learning and TinyML frameworks. Recommended tools include TensorFlow Lite for Microcontrollers, PyTorch Mobile or similar. • No Pre-Trained Global Keywords: Teams cannot use models pre-trained on generic smart-assistant keywords like 'Hey Google' or 'Alexa'. They need to train on a custom key word.

Expected Solution Teams are expected to deliver a robust, deployable system architecture. A successful submission must strictly satisfy the following technical boundaries:

• Hardware & Runtime Environment: The edge software application must run smoothly within an environment restricted to less than 256KB of RAM and consume under 10% CPU utilization while idling in continuous listening mode. Heavy or uncompressed pre-trained transformers are disqualified. Solutions will be formally evaluated on physical low-power microcontrollers (e.g., Raspberry Pi or ESP32). • Model should work for the given custom key word.

How contested this one is

as of 28 Sept
88ideas submitted+24 in 2 days

That puts it 119th of the 240 statements that have any ideas at all, out of 240 on the board. It is moving, so the field here is already forming.

See what the whole field is picking →

Counted from the official portal twice a day. The portal itself only shows today.

What a jury will ask about this

  1. 01“Who actually faces this problem today?”

    What works: Naming one real person and what they do instead right now. Reading the statement back is not an answer, they already read it.

  2. 02“This already exists. Why yours?”

    What works: That existing tools are consumer products. Yours is built for the ministry, works offline, in the local language, on official data.

  3. 03“Then why has nobody solved it yet?”

    What works: The real blocker. No connectivity, no incentive, nobody owns the data. You only know this if you read the ministry's own reports.

All 18 questions, with the trap answers →

More in Smart Automation

See all →