Skip to content

react-native-voice-activator: LLM/AI Assistant Context ​

This page exists to orient AI assistants and developers in one read. Start here, then follow the linked docs for depth.

Full README: README.md | All docs: Documentation table


What This Package Does ​

On-device wake word detection and managed multi-turn voice conversation sessions for React Native and Expo. Say a trigger phrase and the package drives the full loop: listen for speech, transcribe it, call your AI handler, speak the response, then listen again. No cloud required for wake word detection. Speech-to-text and text-to-speech are opt-in, provider-injected, and run on-device.

Supports: React Native 0.86+ · Expo SDK 57+ · iOS 13+ · Android API 26+. Expo Go is NOT supported -- use expo prebuild or EAS Build.


Architecture ​

Public API          src/public/          Singleton functions + React hooks
Orchestration       src/runtime/         Session loop manager (wake→STT→AI→TTS)
Engines             src/engines/         Native-managed engine runtime (Sherpa-ONNX)
Providers           src/providers/       Opt-in STT / TTS / VAD / verification adapters
Internal            src/internal/        Native bridge, event emitters, runtime store
Native Interface    NativeVoiceActivator.ts  TurboModule codegen spec (RN Codegen)

Key design facts:

  • Singleton state. src/public/voice-activator.ts uses module-level variables, not a class. One runtime per app process. initialize() / dispose() manage the lifecycle.
  • Two event emitters. runtimeEvents fires wake word lifecycle events (wakeWordDetected, stateChanged, error, interruption, audioRouteChanged). sessionEvents fires conversation turn events (sessionStarted, sessionListening, sessionTranscribed, sessionSpeaking, sessionTurnComplete, sessionEnded, sessionError). Subscribe with addWakeWordListener and addSessionListener respectively.
  • Provider injection. STT and TTS are passed to initialize() via sttProvider / ttsProvider. The package never owns transcription or synthesis -- it calls your provider at the right moment in the loop.
  • Generation IDs. A counter increments on each new session. Async provider callbacks capture the generation at call time and no-op if it no longer matches. This prevents a slow STT response from a stale session completing into a new one.
  • Barge-in fast-path. If a wake word fires while TTS is speaking, a dedicated path calls ttsProvider.stop() and re-enters the listen stage without waiting for the normal orchestration queue. A wake word arriving during the listening or transcribing stage aborts the capture, discards the partial utterance, and restarts the listening turn. Interruption latency is not yet measured on physical devices.

Key Concepts ​

Session loop. Once sttProvider and ttsProvider are both passed to initialize(), each wake word automatically triggers: sttProvider.transcribe() → aiHandler(transcript) → ttsProvider.speak(response) → optional re-listen. Without both providers, wake word events fire but no session loop activates.

VAD gate. The optional SileroVADEngine sits between wake word detection and STT handoff. It requires sustained speech energy before forwarding to transcription, filtering accidental trigger events.

Speaker verification. The optional SherpaOnnxSpeakerVerificationAdapter compares the speaker's live voiceprint to an enrolled embedding before proceeding with a session turn. Enrollment data is biometric -- see Privacy & Compliance below.

Noise suppression. The optional SherpaOnnxNoiseSuppressionAdapter preprocesses audio before transcription to improve STT accuracy in noisy environments.

Anti-spoofing. The antiSpoofingProvider option gates detection on a liveness score: a score above spoofingThreshold fails the gate. No bundled implementation exists — the native detectSpoofing bridge is a stub returning a constant 0.0, so SherpaOnnxAntiSpoofingAdapter throws on construction rather than passing every input. Supply your own provider to use this.

reListenMode. Set to 'auto' in VoiceSessionConfig and the loop re-enters the listen stage automatically after TTS finishes. Set to 'manual' to wait for an explicit session.listen() call.


Complete API Surface ​

Lifecycle Functions ​

ExportDescription
initialize(options)Load the wake word engine and register providers. Must be called before startDetection().
startDetection()Begin listening for wake words. Engine must be initialized.
stopDetection()Stop listening. Keeps engine loaded -- cheaper to restart than re-initialize.
dispose()Release all native resources. Call on unmount / app background.
getStatus()Returns WakeWordStatus snapshot: state, lastError, isListening, etc.
getSession()Returns the active VoiceSession object, or null if no session is running.
setAudioRoute(route)Switch output between 'default', 'speaker', 'earpiece', and 'bluetooth'.

Events ​

ExportDescription
addWakeWordListener(event, cb)Subscribe to runtime events. Returns { remove() }.
addSessionListener(event, cb)Subscribe to session turn events. Returns { remove() }.

Runtime event names: 'wakeWordDetected' · 'stateChanged' · 'error' · 'interruption' · 'audioRouteChanged'

Session event names: 'sessionStarted' · 'sessionListening' · 'sessionTranscribed' · 'sessionSpeaking' · 'sessionTurnComplete' · 'sessionEnded' · 'sessionError'

Speaker Enrollment ​

ExportDescription
enrollSpeaker(userId, audioBuffer)Extract voiceprint from audio buffer and store in memory. Requires explicit user consent first (GDPR/BIPA).
exportEnrollment()Serialize enrollment to EnrollmentData for persistent storage. Encrypt at rest in your app.
importEnrollment(data)Restore a previously exported enrollment.
clearEnrollment()Remove all biometric data from memory immediately.

React Hooks ​

ExportDescription
useWakeWord()Reactive UseWakeWordResult snapshot: runtime state, last detected phrase, STT and TTS states.
useVoiceSession()Reactive UseVoiceSessionResult snapshot: session state, last transcript, current turn.

Built-In Provider Adapters (opt-in) ​

ExportPeer DependencyDescription
WhisperRNSTTAdapterwhisper.rnOn-device STT via Whisper models.
SherpaOnnxTTSAdapternone (bundled)On-device TTS via Piper VITS models through the bundled Sherpa-ONNX layer.
CustomTTSAdapteronnxruntime-react-nativeOn-device TTS via any Piper ONNX model. Note: may cause duplicate ONNX symbols on iOS -- see iOS ONNX Conflict Resolution.
SileroVADEnginenone (bundled)VAD gate between wake word and STT.
SherpaOnnxSpeakerVerificationAdapternone (bundled)Voiceprint-based speaker identity check.
SherpaOnnxNoiseSuppressionAdapternone (bundled)Audio noise suppression preprocessing.
SherpaOnnxAntiSpoofingAdaptern/aNot implemented — throws on construction. The native bridge is a stub. Supply your own AntiSpoofingProvider instead.

Key Types ​

typescript
// Models are downloaded on demand; call this once before initialize().
prepareModels(options?: ModelPreparationOptions): Promise<ModelPreparationResult>
getModelStatus(): Promise<ModelBundleStatus>

ModelPreparationOptions {
  baseUrl?: string            // defaults to this package's GitHub release
  flatAssets?: boolean        // default true (flattened release asset names)
  force?: boolean             // re-download even if already valid
  onProgress?: (p: ModelPreparationProgress) => void
}

ModelBundleStatus {
  ready: boolean
  directory: string
  bundleVersion: string
  missing: string[]           // manifest-relative paths absent or unverified
  bytesTotal: number
}

// initialize() options
WakeWordInitializationOptions {
  engineConfig?: WakeWordEngineConfiguration
  sttProvider?: SpeechToTextProvider
  ttsProvider?: TextToSpeechProvider
  wakePhrase?: string | string[]      // any English phrase; no training needed
  autoSpeak?: boolean                 // default false; single-shot flow only
  providerTimeoutMs?: number          // default 30000; 0 disables
  session?: VoiceSessionConfig
  speakerVerificationProvider?: SpeakerVerificationProvider
  speakerModelPath?: string           // Android-only, required for the Sherpa adapter
  audioPreprocessingProvider?: AudioPreprocessingProvider
  antiSpoofingProvider?: AntiSpoofingProvider   // no bundled impl; supply your own
  spoofingThreshold?: number          // default 0.5
  verificationThreshold?: number      // default 0.55
  verificationFailureBehavior?: 'open' | 'closed' | 'emit'   // default 'closed'
  vadGateEnabled?: boolean            // default false
  vadGateThreshold?: number           // default 0.5
}

// session loop config
VoiceSessionConfig {
  aiHandler: AIHandler              // (transcript: string) => Promise<string>
  reListenMode: 'auto' | 'manual'
  silenceTimeoutMs?: number         // user never speaks; does NOT cover a hung provider
  maxTurns?: number
  vad?: VADConfig
  providerTimeoutMs?: number        // default 30000; bounds transcribe() and speak()
  aiHandlerTimeoutMs?: number       // default 60000; bounds aiHandler()
}

// provider interfaces (implement these for custom providers)
SpeechToTextProvider {
  name: string
  transcribe(): Promise<TranscriptionResult>
  cancel(): Promise<void>
}
TextToSpeechProvider {
  name: string
  speak(text: string, options?: TTSOptions): Promise<void>
  stop(): Promise<void>
}

// status snapshot
WakeWordStatus {
  state: WakeWordState              // 'idle' | 'initializing' | 'ready' | 'starting' | 'running' | 'interrupted' | 'stopping' | 'stopped' | 'error' | 'unsupported'
  isAvailable: boolean
  isListening: boolean
  canStart: boolean
  reason?: string
  lastError?: WakeWordError | null
}

// error shape
WakeWordError {
  category: WakeWordErrorCategory   // see Error Reference below
  message: string
  recoverable: boolean
}

Minimal Setup ​

Wake Word Only (bare React Native) ​

Validates the runtime without STT or TTS. Say "Hello World" to confirm detection.

typescript
import {
  initialize,
  startDetection,
  stopDetection,
  addWakeWordListener,
  getStatus,
  dispose,
} from 'react-native-voice-activator';

async function runQuickstart() {
  const status = getStatus();
  if (status.state === 'unsupported') {
    console.log('Unsupported:', status.lastError?.message);
    return;
  }

  const sub = addWakeWordListener('wakeWordDetected', (e) => {
    console.log('Detected:', e.detectedPhrase);
  });

  try {
    await initialize();
    await startDetection();
    // say "Hello World"
    await stopDetection();
    await dispose();
  } finally {
    sub.remove();
  }
}

Wake Word Only (Expo) ​

Same API. Add the plugin to app.json first:

json
{
  "expo": {
    "plugins": [
      ["react-native-voice-activator", { "microphonePermissionText": "Microphone is used for wake word detection." }]
    ]
  }
}

Then run npx expo prebuild && cd ios && pod install. Runtime code is identical to bare React Native above.

Full Session (STT + TTS + AI handler) ​

typescript
import {
  initialize,
  startDetection,
  useVoiceSession,
} from 'react-native-voice-activator';
import { WhisperRNSTTAdapter } from 'react-native-voice-activator';
import { SherpaOnnxTTSAdapter } from 'react-native-voice-activator';

// 1. Wire providers + AI handler
await initialize({
  sttProvider: new WhisperRNSTTAdapter({ modelId: 'whisper-tiny-en' }),
  ttsProvider: new SherpaOnnxTTSAdapter({
    modelPath: `${RNFS.DocumentDirectoryPath}/sherpa-tts/en_US-ryan-low.onnx`,
    tokensPath: `${RNFS.DocumentDirectoryPath}/sherpa-tts/tokens.txt`,
    dataDir:    `${RNFS.DocumentDirectoryPath}/sherpa-tts/espeak-ng-data`,
  }),
  session: {
    aiHandler: async (transcript) => {
      // call your LLM or backend here
      return `You said: ${transcript}`;
    },
    reListenMode: 'auto',
  },
});

await startDetection();

// 2. React component subscribes to session state
function ConversationUI() {
  const { state, lastTranscript } = useVoiceSession();
  return <Text>{state === 'listening' ? 'Listening...' : lastTranscript}</Text>;
}

See Conversation Session for the full session event reference.


Common Patterns ​

Pattern 1: Wake word only, no session ​

Use when you want to react to a trigger phrase but drive STT/TTS yourself.

typescript
await initialize(); // no sttProvider or ttsProvider

addWakeWordListener('wakeWordDetected', async (event) => {
  console.log('Triggered by:', event.detectedPhrase);
  // your own downstream logic here
});

await startDetection();

Pattern 2: Custom provider (implement your own STT) ​

Implement the SpeechToTextProvider interface and pass it to initialize().

typescript
import type { SpeechToTextProvider, TranscriptionResult } from 'react-native-voice-activator';

class MySTTProvider implements SpeechToTextProvider {
  readonly name = 'MySTT';

  async transcribe(): Promise<TranscriptionResult> {
    // start recording, stop on silence, return transcript
    return { text: 'hello world', confidence: 0.95, provider: 'MySTT' };
  }

  async cancel(): Promise<void> {
    // stop recording and discard
  }
}

await initialize({
  sttProvider: new MySTTProvider(),
  ttsProvider: myTTSProvider,
  session: { aiHandler: myAIHandler, reListenMode: 'auto' },
});

Pattern 3: Manual re-listen (you control when to listen again) ​

typescript
await initialize({
  sttProvider: stt,
  ttsProvider: tts,
  session: {
    reListenMode: 'manual',
    aiHandler: async (transcript) => {
      const response = await callMyBackend(transcript);
      return response;
    },
  },
});

// Later, when ready for the next turn:
const session = getSession();
await session?.listen();
typescript
import {
  enrollSpeaker,
  exportEnrollment,
  importEnrollment,
  clearEnrollment,
} from 'react-native-voice-activator';

// Enroll (requires explicit user consent first -- GDPR/BIPA requirement)
async function enrollWithConsent(userId: string, audioBuffer: ArrayBuffer) {
  const hasConsent = await showConsentDialog();
  if (!hasConsent) throw new Error('User consent required');

  await enrollSpeaker(userId, audioBuffer);
  const data = await exportEnrollment();
  await secureStorage.save(`enrollment_${userId}`, JSON.stringify(data));
}

// Restore on app launch
async function restoreEnrollment(userId: string) {
  const raw = await secureStorage.get(`enrollment_${userId}`);
  if (raw) await importEnrollment(JSON.parse(raw));
}

// Delete on account deletion or consent revocation
async function deleteEnrollment(userId: string) {
  await clearEnrollment();
  await secureStorage.delete(`enrollment_${userId}`);
}

Error Reference ​

WakeWordErrorCategoryCauseRecovery
permissionMicrophone permission deniedRequest RECORD_AUDIO (Android) or check NSMicrophoneUsageDescription (iOS)
lifecycleAPI called in wrong state (e.g. startDetection before initialize)Check getStatus().state before calling lifecycle functions
configurationInvalid or missing initialization optionsReview WakeWordInitializationOptions -- both sttProvider and ttsProvider are required for sessions
engineNative engine failed to load or crashedCheck that model assets are bundled correctly; see Getting Started
platformOS-level constraint (background mode, foreground service), or the native module is absent (runtime_unavailable, not recoverable)See Background Behavior
configuration / models_not_preparedThe on-demand model bundle is absent or unverified (not recoverable)Call prepareModels(), or pass engineConfig.assetKeys.modelAssetKey
configuration / wake_phrase_invalidA wakePhrase failed validation (not recoverable)Message lists every problem; see the phrase rules
configuration / wake_phrase_conflictBoth wakePhrase and keywordAssetKey were supplied (not recoverable)Pick one
configuration / wake_phrase_unsupported_rootwakePhrase used with an app-bundled modelAssetKey (not recoverable)Generate a keywords file offline, pass keywordAssetKey
internalUnexpected runtime errorFile a bug; include getStatus().lastError.message

All errors carry recoverable: boolean, answering only: can the same call with the same options succeed?

  • recoverable: true — retry may work. Transient provider failures and timeouts (stt_timeout, tts_timeout, ai_handler_timeout), most permission and lifecycle errors.
  • recoverable: false — retry cannot work. Every configuration error (fix the options, then call initialize() again) and runtime_unavailable (the native module is missing from the build; rebuild with pod install / expo prebuild — re-initialising will not help).

Privacy & Compliance Summary ​

enrollSpeaker() processes biometric voiceprint data. Your app must:

  • Obtain explicit consent before calling enrollSpeaker() (required by GDPR Art. 9, BIPA, CCPA)
  • Encrypt exportEnrollment() output at rest
  • Implement deletion: clearEnrollment() + delete stored data
  • Disclose biometric data collection in your privacy policy

All processing is on-device. No data leaves the device through this library.


Documentation ​

DocumentDescription
LLM/AI Assistant ContextThis page -- one-stop orientation
Getting StartedStep-by-step setup from zero to first detection
Bare React Native SetupNative project configuration for bare React Native
Expo SetupConfig plugin and prebuild setup for Expo
Conversation SessionFull session API, events, barge-in, and React hooks
Background BehavioriOS and Android background detection constraints
WhisperRN STT ProviderOn-device speech-to-text setup
Custom TTS ProviderOn-device text-to-speech setup
iOS ONNX Conflict ResolutionFix duplicate ONNX symbol linker errors
Wake Word TrainingTrain a custom wake word offline
TTS Voice CloningTrain a custom TTS voice
TroubleshootingFull error category reference
Android TTS SetupAndroid model asset placement for SherpaOnnxTTSAdapter
MigrationVersion migration guide
App Store SubmissionApp Store review guidance