react-native-voice-activator: LLM/AI Assistant Context
This page exists to orient AI assistants and developers in one read. Start here, then follow the linked docs for depth.
Full README: README.md | All docs: Documentation table
What This Package Does
On-device wake word detection and managed multi-turn voice conversation sessions for React Native and Expo. Say a trigger phrase and the package drives the full loop: listen for speech, transcribe it, call your AI handler, speak the response, then listen again. No cloud required for wake word detection. Speech-to-text and text-to-speech are opt-in, provider-injected, and run on-device.
Supports: React Native 0.86+ · Expo SDK 57+ · iOS 13+ · Android API 26+. Expo Go is NOT supported -- use expo prebuild or EAS Build.
Architecture
Public API src/public/ Singleton functions + React hooks
Orchestration src/runtime/ Session loop manager (wake→STT→AI→TTS)
Engines src/engines/ Native-managed engine runtime (Sherpa-ONNX)
Providers src/providers/ Opt-in STT / TTS / VAD / verification adapters
Internal src/internal/ Native bridge, event emitters, runtime store
Native Interface NativeVoiceActivator.ts TurboModule codegen spec (RN Codegen)Key design facts:
- Singleton state.
src/public/voice-activator.tsuses module-level variables, not a class. One runtime per app process.initialize()/dispose()manage the lifecycle. - Two event emitters.
runtimeEventsfires wake word lifecycle events (wakeWordDetected,stateChanged,error,interruption,audioRouteChanged).sessionEventsfires conversation turn events (sessionStarted,sessionListening,sessionTranscribed,sessionSpeaking,sessionTurnComplete,sessionEnded,sessionError). Subscribe withaddWakeWordListenerandaddSessionListenerrespectively. - Provider injection. STT and TTS are passed to
initialize()viasttProvider/ttsProvider. The package never owns transcription or synthesis -- it calls your provider at the right moment in the loop. - Generation IDs. A counter increments on each new session. Async provider callbacks capture the generation at call time and no-op if it no longer matches. This prevents a slow STT response from a stale session completing into a new one.
- Barge-in fast-path. If a wake word fires while TTS is speaking, a dedicated path calls
ttsProvider.stop()and re-enters the listen stage without waiting for the normal orchestration queue. A wake word arriving during thelisteningortranscribingstage aborts the capture, discards the partial utterance, and restarts the listening turn. Interruption latency is not yet measured on physical devices.
Key Concepts
Session loop. Once sttProvider and ttsProvider are both passed to initialize(), each wake word automatically triggers: sttProvider.transcribe() → aiHandler(transcript) → ttsProvider.speak(response) → optional re-listen. Without both providers, wake word events fire but no session loop activates.
VAD gate. The optional SileroVADEngine sits between wake word detection and STT handoff. It requires sustained speech energy before forwarding to transcription, filtering accidental trigger events.
Speaker verification. The optional SherpaOnnxSpeakerVerificationAdapter compares the speaker's live voiceprint to an enrolled embedding before proceeding with a session turn. Enrollment data is biometric -- see Privacy & Compliance below.
Noise suppression. The optional SherpaOnnxNoiseSuppressionAdapter preprocesses audio before transcription to improve STT accuracy in noisy environments.
Anti-spoofing. The antiSpoofingProvider option gates detection on a liveness score: a score above spoofingThreshold fails the gate. No bundled implementation exists — the native detectSpoofing bridge is a stub returning a constant 0.0, so SherpaOnnxAntiSpoofingAdapter throws on construction rather than passing every input. Supply your own provider to use this.
reListenMode. Set to 'auto' in VoiceSessionConfig and the loop re-enters the listen stage automatically after TTS finishes. Set to 'manual' to wait for an explicit session.listen() call.
Complete API Surface
Lifecycle Functions
| Export | Description |
|---|---|
initialize(options) | Load the wake word engine and register providers. Must be called before startDetection(). |
startDetection() | Begin listening for wake words. Engine must be initialized. |
stopDetection() | Stop listening. Keeps engine loaded -- cheaper to restart than re-initialize. |
dispose() | Release all native resources. Call on unmount / app background. |
getStatus() | Returns WakeWordStatus snapshot: state, lastError, isListening, etc. |
getSession() | Returns the active VoiceSession object, or null if no session is running. |
setAudioRoute(route) | Switch output between 'default', 'speaker', 'earpiece', and 'bluetooth'. |
Events
| Export | Description |
|---|---|
addWakeWordListener(event, cb) | Subscribe to runtime events. Returns { remove() }. |
addSessionListener(event, cb) | Subscribe to session turn events. Returns { remove() }. |
Runtime event names: 'wakeWordDetected' · 'stateChanged' · 'error' · 'interruption' · 'audioRouteChanged'
Session event names: 'sessionStarted' · 'sessionListening' · 'sessionTranscribed' · 'sessionSpeaking' · 'sessionTurnComplete' · 'sessionEnded' · 'sessionError'
Speaker Enrollment
| Export | Description |
|---|---|
enrollSpeaker(userId, audioBuffer) | Extract voiceprint from audio buffer and store in memory. Requires explicit user consent first (GDPR/BIPA). |
exportEnrollment() | Serialize enrollment to EnrollmentData for persistent storage. Encrypt at rest in your app. |
importEnrollment(data) | Restore a previously exported enrollment. |
clearEnrollment() | Remove all biometric data from memory immediately. |
React Hooks
| Export | Description |
|---|---|
useWakeWord() | Reactive UseWakeWordResult snapshot: runtime state, last detected phrase, STT and TTS states. |
useVoiceSession() | Reactive UseVoiceSessionResult snapshot: session state, last transcript, current turn. |
Built-In Provider Adapters (opt-in)
| Export | Peer Dependency | Description |
|---|---|---|
WhisperRNSTTAdapter | whisper.rn | On-device STT via Whisper models. |
SherpaOnnxTTSAdapter | none (bundled) | On-device TTS via Piper VITS models through the bundled Sherpa-ONNX layer. |
CustomTTSAdapter | onnxruntime-react-native | On-device TTS via any Piper ONNX model. Note: may cause duplicate ONNX symbols on iOS -- see iOS ONNX Conflict Resolution. |
SileroVADEngine | none (bundled) | VAD gate between wake word and STT. |
SherpaOnnxSpeakerVerificationAdapter | none (bundled) | Voiceprint-based speaker identity check. |
SherpaOnnxNoiseSuppressionAdapter | none (bundled) | Audio noise suppression preprocessing. |
SherpaOnnxAntiSpoofingAdapter | n/a | Not implemented — throws on construction. The native bridge is a stub. Supply your own AntiSpoofingProvider instead. |
Key Types
// Models are downloaded on demand; call this once before initialize().
prepareModels(options?: ModelPreparationOptions): Promise<ModelPreparationResult>
getModelStatus(): Promise<ModelBundleStatus>
ModelPreparationOptions {
baseUrl?: string // defaults to this package's GitHub release
flatAssets?: boolean // default true (flattened release asset names)
force?: boolean // re-download even if already valid
onProgress?: (p: ModelPreparationProgress) => void
}
ModelBundleStatus {
ready: boolean
directory: string
bundleVersion: string
missing: string[] // manifest-relative paths absent or unverified
bytesTotal: number
}
// initialize() options
WakeWordInitializationOptions {
engineConfig?: WakeWordEngineConfiguration
sttProvider?: SpeechToTextProvider
ttsProvider?: TextToSpeechProvider
wakePhrase?: string | string[] // any English phrase; no training needed
autoSpeak?: boolean // default false; single-shot flow only
providerTimeoutMs?: number // default 30000; 0 disables
session?: VoiceSessionConfig
speakerVerificationProvider?: SpeakerVerificationProvider
speakerModelPath?: string // Android-only, required for the Sherpa adapter
audioPreprocessingProvider?: AudioPreprocessingProvider
antiSpoofingProvider?: AntiSpoofingProvider // no bundled impl; supply your own
spoofingThreshold?: number // default 0.5
verificationThreshold?: number // default 0.55
verificationFailureBehavior?: 'open' | 'closed' | 'emit' // default 'closed'
vadGateEnabled?: boolean // default false
vadGateThreshold?: number // default 0.5
}
// session loop config
VoiceSessionConfig {
aiHandler: AIHandler // (transcript: string) => Promise<string>
reListenMode: 'auto' | 'manual'
silenceTimeoutMs?: number // user never speaks; does NOT cover a hung provider
maxTurns?: number
vad?: VADConfig
providerTimeoutMs?: number // default 30000; bounds transcribe() and speak()
aiHandlerTimeoutMs?: number // default 60000; bounds aiHandler()
}
// provider interfaces (implement these for custom providers)
SpeechToTextProvider {
name: string
transcribe(): Promise<TranscriptionResult>
cancel(): Promise<void>
}
TextToSpeechProvider {
name: string
speak(text: string, options?: TTSOptions): Promise<void>
stop(): Promise<void>
}
// status snapshot
WakeWordStatus {
state: WakeWordState // 'idle' | 'initializing' | 'ready' | 'starting' | 'running' | 'interrupted' | 'stopping' | 'stopped' | 'error' | 'unsupported'
isAvailable: boolean
isListening: boolean
canStart: boolean
reason?: string
lastError?: WakeWordError | null
}
// error shape
WakeWordError {
category: WakeWordErrorCategory // see Error Reference below
message: string
recoverable: boolean
}Minimal Setup
Wake Word Only (bare React Native)
Validates the runtime without STT or TTS. Say "Hello World" to confirm detection.
import {
initialize,
startDetection,
stopDetection,
addWakeWordListener,
getStatus,
dispose,
} from 'react-native-voice-activator';
async function runQuickstart() {
const status = getStatus();
if (status.state === 'unsupported') {
console.log('Unsupported:', status.lastError?.message);
return;
}
const sub = addWakeWordListener('wakeWordDetected', (e) => {
console.log('Detected:', e.detectedPhrase);
});
try {
await initialize();
await startDetection();
// say "Hello World"
await stopDetection();
await dispose();
} finally {
sub.remove();
}
}Wake Word Only (Expo)
Same API. Add the plugin to app.json first:
{
"expo": {
"plugins": [
["react-native-voice-activator", { "microphonePermissionText": "Microphone is used for wake word detection." }]
]
}
}Then run npx expo prebuild && cd ios && pod install. Runtime code is identical to bare React Native above.
Full Session (STT + TTS + AI handler)
import {
initialize,
startDetection,
useVoiceSession,
} from 'react-native-voice-activator';
import { WhisperRNSTTAdapter } from 'react-native-voice-activator';
import { SherpaOnnxTTSAdapter } from 'react-native-voice-activator';
// 1. Wire providers + AI handler
await initialize({
sttProvider: new WhisperRNSTTAdapter({ modelId: 'whisper-tiny-en' }),
ttsProvider: new SherpaOnnxTTSAdapter({
modelPath: `${RNFS.DocumentDirectoryPath}/sherpa-tts/en_US-ryan-low.onnx`,
tokensPath: `${RNFS.DocumentDirectoryPath}/sherpa-tts/tokens.txt`,
dataDir: `${RNFS.DocumentDirectoryPath}/sherpa-tts/espeak-ng-data`,
}),
session: {
aiHandler: async (transcript) => {
// call your LLM or backend here
return `You said: ${transcript}`;
},
reListenMode: 'auto',
},
});
await startDetection();
// 2. React component subscribes to session state
function ConversationUI() {
const { state, lastTranscript } = useVoiceSession();
return <Text>{state === 'listening' ? 'Listening...' : lastTranscript}</Text>;
}See Conversation Session for the full session event reference.
Common Patterns
Pattern 1: Wake word only, no session
Use when you want to react to a trigger phrase but drive STT/TTS yourself.
await initialize(); // no sttProvider or ttsProvider
addWakeWordListener('wakeWordDetected', async (event) => {
console.log('Triggered by:', event.detectedPhrase);
// your own downstream logic here
});
await startDetection();Pattern 2: Custom provider (implement your own STT)
Implement the SpeechToTextProvider interface and pass it to initialize().
import type { SpeechToTextProvider, TranscriptionResult } from 'react-native-voice-activator';
class MySTTProvider implements SpeechToTextProvider {
readonly name = 'MySTT';
async transcribe(): Promise<TranscriptionResult> {
// start recording, stop on silence, return transcript
return { text: 'hello world', confidence: 0.95, provider: 'MySTT' };
}
async cancel(): Promise<void> {
// stop recording and discard
}
}
await initialize({
sttProvider: new MySTTProvider(),
ttsProvider: myTTSProvider,
session: { aiHandler: myAIHandler, reListenMode: 'auto' },
});Pattern 3: Manual re-listen (you control when to listen again)
await initialize({
sttProvider: stt,
ttsProvider: tts,
session: {
reListenMode: 'manual',
aiHandler: async (transcript) => {
const response = await callMyBackend(transcript);
return response;
},
},
});
// Later, when ready for the next turn:
const session = getSession();
await session?.listen();Pattern 4: Speaker enrollment with consent gate
import {
enrollSpeaker,
exportEnrollment,
importEnrollment,
clearEnrollment,
} from 'react-native-voice-activator';
// Enroll (requires explicit user consent first -- GDPR/BIPA requirement)
async function enrollWithConsent(userId: string, audioBuffer: ArrayBuffer) {
const hasConsent = await showConsentDialog();
if (!hasConsent) throw new Error('User consent required');
await enrollSpeaker(userId, audioBuffer);
const data = await exportEnrollment();
await secureStorage.save(`enrollment_${userId}`, JSON.stringify(data));
}
// Restore on app launch
async function restoreEnrollment(userId: string) {
const raw = await secureStorage.get(`enrollment_${userId}`);
if (raw) await importEnrollment(JSON.parse(raw));
}
// Delete on account deletion or consent revocation
async function deleteEnrollment(userId: string) {
await clearEnrollment();
await secureStorage.delete(`enrollment_${userId}`);
}Error Reference
WakeWordErrorCategory | Cause | Recovery |
|---|---|---|
permission | Microphone permission denied | Request RECORD_AUDIO (Android) or check NSMicrophoneUsageDescription (iOS) |
lifecycle | API called in wrong state (e.g. startDetection before initialize) | Check getStatus().state before calling lifecycle functions |
configuration | Invalid or missing initialization options | Review WakeWordInitializationOptions -- both sttProvider and ttsProvider are required for sessions |
engine | Native engine failed to load or crashed | Check that model assets are bundled correctly; see Getting Started |
platform | OS-level constraint (background mode, foreground service), or the native module is absent (runtime_unavailable, not recoverable) | See Background Behavior |
configuration / models_not_prepared | The on-demand model bundle is absent or unverified (not recoverable) | Call prepareModels(), or pass engineConfig.assetKeys.modelAssetKey |
configuration / wake_phrase_invalid | A wakePhrase failed validation (not recoverable) | Message lists every problem; see the phrase rules |
configuration / wake_phrase_conflict | Both wakePhrase and keywordAssetKey were supplied (not recoverable) | Pick one |
configuration / wake_phrase_unsupported_root | wakePhrase used with an app-bundled modelAssetKey (not recoverable) | Generate a keywords file offline, pass keywordAssetKey |
internal | Unexpected runtime error | File a bug; include getStatus().lastError.message |
All errors carry recoverable: boolean, answering only: can the same call with the same options succeed?
recoverable: true— retry may work. Transient provider failures and timeouts (stt_timeout,tts_timeout,ai_handler_timeout), mostpermissionandlifecycleerrors.recoverable: false— retry cannot work. Everyconfigurationerror (fix the options, then callinitialize()again) andruntime_unavailable(the native module is missing from the build; rebuild withpod install/expo prebuild— re-initialising will not help).
Privacy & Compliance Summary
enrollSpeaker() processes biometric voiceprint data. Your app must:
- Obtain explicit consent before calling
enrollSpeaker()(required by GDPR Art. 9, BIPA, CCPA) - Encrypt
exportEnrollment()output at rest - Implement deletion:
clearEnrollment()+ delete stored data - Disclose biometric data collection in your privacy policy
All processing is on-device. No data leaves the device through this library.
Documentation
| Document | Description |
|---|---|
| LLM/AI Assistant Context | This page -- one-stop orientation |
| Getting Started | Step-by-step setup from zero to first detection |
| Bare React Native Setup | Native project configuration for bare React Native |
| Expo Setup | Config plugin and prebuild setup for Expo |
| Conversation Session | Full session API, events, barge-in, and React hooks |
| Background Behavior | iOS and Android background detection constraints |
| WhisperRN STT Provider | On-device speech-to-text setup |
| Custom TTS Provider | On-device text-to-speech setup |
| iOS ONNX Conflict Resolution | Fix duplicate ONNX symbol linker errors |
| Wake Word Training | Train a custom wake word offline |
| TTS Voice Cloning | Train a custom TTS voice |
| Troubleshooting | Full error category reference |
| Android TTS Setup | Android model asset placement for SherpaOnnxTTSAdapter |
| Migration | Version migration guide |
| App Store Submission | App Store review guidance |