IEEE TSP/TASLP Submission Package — NanoMamba

Submission 방향

타깃: IEEE TSP / IEEE TASLP 계열

핵심 주제: NanoMamba 기반 초경량·저지연·저전력 음성/신호처리 AI 엔진

Working Title 후보

  1. NanoMamba: An Ultra-Lightweight State-Space Model for Robust Keyword Spotting on Edge Devices
  2. Noise-Robust On-Device Keyword Spotting with Sub-5KB Mamba-Based Architecture
  3. Ultra-Low-Power Speech Command Recognition under Industrial Noise via Adaptive State-Space Modeling

Core Contribution

  • 초경량 모델 크기: 약 4.8KB 수준의 KWS / speech command recognition 모델
  • 초저지연 추론: Cortex-M 계열 MCU에서 약 0.94ms 응답
  • 초저전력 운용: CR2032 배터리 기반 장기 구동 가능성
  • 산업 소음 강건성: -15dB 수준의 극한 소음 환경에서도 robust keyword spotting 지향
  • 하드웨어 추가 없음: 별도 noise-cancellation chip 없이 모델 구조 자체로 노이즈 적응
  • Edge-first 설계: ARM Cortex-M, ESP32-S3, nRF5340 등 초소형 디바이스 배포 가능

논문 포지셔닝

Problem

기존 KWS / speech command recognition 모델은 조용한 환경에서는 높은 정확도를 보이지만, 공장·설비·차량 등 실제 산업 소음 환경에서는 정확도가 급격히 하락한다. 또한 MCU급 엣지 디바이스에서는 모델 크기, 연산량, latency, 전력 소모가 동시에 제한된다.

Proposed Method

NanoMamba는 SA-SSM 기반의 초경량 state-space architecture와 DualPCEN 기반 noise-robust preprocessing, selective scan + bypass 구조를 결합해, 작은 파라미터 수에서도 소음 구간과 음성 구간의 temporal dynamics를 효율적으로 모델링한다.

Key Claim

NanoMamba는 CNN/RNN 계열 baseline 대비 훨씬 낮은 MACs와 latency로, 산업 소음 환경에서 robust keyword spotting을 수행하는 초소형 on-device speech AI 엔진이다.

실험 설계

Dataset

  • Google Speech Commands
  • 산업 소음 / 공장 소음 augmentation dataset
  • 자체 수집 field noise dataset
  • -5dB, -10dB, -15dB SNR 조건별 평가

Baselines

  • DS-CNN
  • TC-ResNet
  • BC-ResNet-1
  • CRNN
  • TinyConv
  • 기존 lightweight KWS SSM/Mamba variant

Metrics

  • Accuracy / F1-score
  • False Accept Rate / False Reject Rate
  • MACs
  • Parameter count
  • Model size
  • Latency on MCU
  • Energy per inference
  • Robustness across SNR levels

Figure / Table 구성

  1. Architecture Overview — SA-SSM + DualPCEN + bypass 구조
  2. Noise Robustness Curve — SNR별 accuracy 비교
  3. Latency vs Accuracy Plot — baseline 대비 Pareto frontier
  4. Model Size vs Accuracy Table — 4.8KB 모델의 차별성 강조
  5. Edge Deployment Diagram — ARM Cortex-M / ESP32-S3 / nRF5340 deployment
  6. Industrial Use Case — 공장 음성 명령 인식 시나리오

Abstract 초안

We present NanoMamba, an ultra-lightweight state-space model for robust keyword spotting on resource-constrained edge devices. Existing keyword spotting systems often suffer severe accuracy degradation under industrial noise and require additional denoising modules or computationally expensive neural architectures. NanoMamba addresses this challenge through a compact selective state-space design combined with noise-aware front-end processing, enabling robust temporal modeling with minimal memory and computation. With a model footprint of approximately 4.8KB and sub-millisecond inference latency on microcontroller-class hardware, NanoMamba achieves efficient on-device speech command recognition without cloud connectivity or dedicated noise-cancellation hardware. Experiments under multiple signal-to-noise ratio conditions demonstrate that NanoMamba offers a favorable accuracy-latency-energy trade-off compared with lightweight CNN and recurrent baselines. These results suggest that state-space architectures can provide a practical foundation for always-on, battery-powered speech interfaces in industrial edge AI applications.

국내 IP 우선출원 전략

원칙: 논문 공개·학회 제출·데모 공개 전에 국내 특허 우선출원을 먼저 진행한다.

국내 출원일을 확보한 뒤, 논문 제출 및 해외 PCT/미국/유럽/일본 출원 전략으로 확장한다.

우선출원 대상 발명

  1. NanoMamba 초경량 음성 인식 모델 구조
    • SA-SSM 기반 초소형 KWS architecture
    • selective scan + bypass 구조
    • sub-5KB 모델 압축 및 MCU 배포 구조
  2. DualPCEN 기반 산업 소음 강건 전처리
    • 공장 소음 / 차량 소음 / 설비 소음 환경에서의 adaptive noise-aware preprocessing
    • -15dB SNR 조건 대응 구조
    • 별도 noise-cancellation chip 없이 모델 내부에서 노이즈 적응
  3. 초저전력 Always-on 음성 AI 운용 방법
    • CR2032 배터리 기반 장기 구동
    • event-triggered inference
    • MCU/NPU 기반 저전력 wake-word detection
  4. Edge deployment 및 runtime 최적화
    • ARM Cortex-M / ESP32-S3 / nRF5340 대상 quantization, memory layout, inference scheduling
    • NanoRT 또는 경량 runtime 연동 방식

특허 명세서 제목 후보

  • 산업 소음 환경에서의 초경량 상태공간 모델 기반 키워드 인식 장치 및 방법
  • 엣지 디바이스용 저전력 음성 명령 인식 모델 및 그 운용 방법
  • 노이즈 적응형 전처리와 상태공간 모델을 이용한 온디바이스 음성 인식 방법
  • 마이크로컨트롤러 기반 초소형 키워드 스팟팅 모델의 추론 최적화 방법

출원 전 체크리스트

논문/포스터/웹페이지 공개 전 신규성 훼손 여부 점검
핵심 architecture diagram 특허 도면화
baseline 대비 효과를 “기술적 효과”로 정리
claim scope: 모델 구조 / 전처리 / 추론 방법 / 장치 / 시스템으로 분리
국내 우선출원 후 12개월 내 PCT 또는 해외 개별국 출원 일정 관리
논문 제출본과 특허 명세서 간 공개 범위 정합성 확인

준비해야 할 자료

최종 benchmark table 정리
SNR별 accuracy / F1 결과 정리
Cortex-M 실제 latency 측정 로그
power profiling 자료
architecture diagram 제작
baseline 재현 코드 정리
ablation study: DualPCEN / bypass / SSM block 제거 비교
IEEE template 기반 manuscript 작성
related work 정리: KWS, SSM, Mamba, TinyML, robust speech recognition
국내 특허 우선출원 명세서 초안
국내 특허 청구항 초안

🌐 Translations / 翻译 / 翻訳 / Übersetzungen

‣

🇺🇸 English

‣

🇨🇳 中文

‣

🇯🇵 日本語

‣

🇩🇪 Deutsch