Submission 방향
타깃: IEEE TSP / IEEE TASLP 계열
핵심 주제: NanoMamba 기반 초경량·저지연·저전력 음성/신호처리 AI 엔진
Working Title 후보
- NanoMamba: An Ultra-Lightweight State-Space Model for Robust Keyword Spotting on Edge Devices
- Noise-Robust On-Device Keyword Spotting with Sub-5KB Mamba-Based Architecture
- Ultra-Low-Power Speech Command Recognition under Industrial Noise via Adaptive State-Space Modeling
Core Contribution
- 초경량 모델 크기: 약 4.8KB 수준의 KWS / speech command recognition 모델
- 초저지연 추론: Cortex-M 계열 MCU에서 약 0.94ms 응답
- 초저전력 운용: CR2032 배터리 기반 장기 구동 가능성
- 산업 소음 강건성: -15dB 수준의 극한 소음 환경에서도 robust keyword spotting 지향
- 하드웨어 추가 없음: 별도 noise-cancellation chip 없이 모델 구조 자체로 노이즈 적응
- Edge-first 설계: ARM Cortex-M, ESP32-S3, nRF5340 등 초소형 디바이스 배포 가능
논문 포지셔닝
Problem
기존 KWS / speech command recognition 모델은 조용한 환경에서는 높은 정확도를 보이지만, 공장·설비·차량 등 실제 산업 소음 환경에서는 정확도가 급격히 하락한다. 또한 MCU급 엣지 디바이스에서는 모델 크기, 연산량, latency, 전력 소모가 동시에 제한된다.
Proposed Method
NanoMamba는 SA-SSM 기반의 초경량 state-space architecture와 DualPCEN 기반 noise-robust preprocessing, selective scan + bypass 구조를 결합해, 작은 파라미터 수에서도 소음 구간과 음성 구간의 temporal dynamics를 효율적으로 모델링한다.
Key Claim
NanoMamba는 CNN/RNN 계열 baseline 대비 훨씬 낮은 MACs와 latency로, 산업 소음 환경에서 robust keyword spotting을 수행하는 초소형 on-device speech AI 엔진이다.
실험 설계
Dataset
- Google Speech Commands
- 산업 소음 / 공장 소음 augmentation dataset
- 자체 수집 field noise dataset
- -5dB, -10dB, -15dB SNR 조건별 평가
Baselines
- DS-CNN
- TC-ResNet
- BC-ResNet-1
- CRNN
- TinyConv
- 기존 lightweight KWS SSM/Mamba variant
Metrics
- Accuracy / F1-score
- False Accept Rate / False Reject Rate
- MACs
- Parameter count
- Model size
- Latency on MCU
- Energy per inference
- Robustness across SNR levels
Figure / Table 구성
- Architecture Overview — SA-SSM + DualPCEN + bypass 구조
- Noise Robustness Curve — SNR별 accuracy 비교
- Latency vs Accuracy Plot — baseline 대비 Pareto frontier
- Model Size vs Accuracy Table — 4.8KB 모델의 차별성 강조
- Edge Deployment Diagram — ARM Cortex-M / ESP32-S3 / nRF5340 deployment
- Industrial Use Case — 공장 음성 명령 인식 시나리오
Abstract 초안
We present NanoMamba, an ultra-lightweight state-space model for robust keyword spotting on resource-constrained edge devices. Existing keyword spotting systems often suffer severe accuracy degradation under industrial noise and require additional denoising modules or computationally expensive neural architectures. NanoMamba addresses this challenge through a compact selective state-space design combined with noise-aware front-end processing, enabling robust temporal modeling with minimal memory and computation. With a model footprint of approximately 4.8KB and sub-millisecond inference latency on microcontroller-class hardware, NanoMamba achieves efficient on-device speech command recognition without cloud connectivity or dedicated noise-cancellation hardware. Experiments under multiple signal-to-noise ratio conditions demonstrate that NanoMamba offers a favorable accuracy-latency-energy trade-off compared with lightweight CNN and recurrent baselines. These results suggest that state-space architectures can provide a practical foundation for always-on, battery-powered speech interfaces in industrial edge AI applications.
국내 IP 우선출원 전략
원칙: 논문 공개·학회 제출·데모 공개 전에 국내 특허 우선출원을 먼저 진행한다.
국내 출원일을 확보한 뒤, 논문 제출 및 해외 PCT/미국/유럽/일본 출원 전략으로 확장한다.
우선출원 대상 발명
- NanoMamba 초경량 음성 인식 모델 구조
- SA-SSM 기반 초소형 KWS architecture
- selective scan + bypass 구조
- sub-5KB 모델 압축 및 MCU 배포 구조
- DualPCEN 기반 산업 소음 강건 전처리
- 공장 소음 / 차량 소음 / 설비 소음 환경에서의 adaptive noise-aware preprocessing
- -15dB SNR 조건 대응 구조
- 별도 noise-cancellation chip 없이 모델 내부에서 노이즈 적응
- 초저전력 Always-on 음성 AI 운용 방법
- CR2032 배터리 기반 장기 구동
- event-triggered inference
- MCU/NPU 기반 저전력 wake-word detection
- Edge deployment 및 runtime 최적화
- ARM Cortex-M / ESP32-S3 / nRF5340 대상 quantization, memory layout, inference scheduling
- NanoRT 또는 경량 runtime 연동 방식
특허 명세서 제목 후보
- 산업 소음 환경에서의 초경량 상태공간 모델 기반 키워드 인식 장치 및 방법
- 엣지 디바이스용 저전력 음성 명령 인식 모델 및 그 운용 방법
- 노이즈 적응형 전처리와 상태공간 모델을 이용한 온디바이스 음성 인식 방법
- 마이크로컨트롤러 기반 초소형 키워드 스팟팅 모델의 추론 최적화 방법