Packages
mob_speech
0.1.0
Speech-to-text for Mob apps: Android SpeechRecognizer / iOS SFSpeechRecognizer, pluggable engines, scriptable fake for tests
Current section
Files
Jump to
Current section
Files
mob_speech
README.md
README.md
# mob_speech
Speech-to-text for apps built with [Mob](https://hexdocs.pm/mob), as a plugin.
**Speech-to-text only.** Text-to-speech stays in mob core (`Mob.Speech`).
Engines are pluggable. The default `:platform` engine uses Android's
`SpeechRecognizer` and iOS's `SFSpeechRecognizer` with an `AVAudioEngine`
input tap. `MobSpeech.Engine.Fake` plays a script, so tests and agents can
drive dictation without a microphone. Other packages add engines by
implementing `MobSpeech.Engine`.
## Installation
```elixir
# mix.exs
{:mob_speech, "~> 0.1"}
# mob.exs
config :mob, :plugins, [:mob_speech]
```
Apps generated by `mix mob.new` already trust the first-party signing key in
`config :mob, :trusted_plugins`. To add the entry by hand:
```elixir
config :mob, :trusted_plugins, %{
mob_speech: "ed25519:nc56w+1Kx0gIt/4EkHxnMZCKHMzp4+S5kS/HoSzEZkg="
}
```
Once you activate the plugin, the build merges in:
- **Android:** `android.permission.RECORD_AUDIO`.
- **iOS:** `NSSpeechRecognitionUsageDescription`, which you can override in
your own `Info.plist`, and the `Speech` and `AVFoundation` frameworks.
`NSMicrophoneUsageDescription` must already be in your `Info.plist`. Apps
generated by `mix mob.new` have it; core owns the microphone key, so this
plugin doesn't declare it.
- **Permissions:** the plugin registers the `:speech` capability for
`Mob.Permissions.request(socket, :speech)`. On Android that's
`RECORD_AUDIO`. On iOS it's speech-recognition authorisation plus the
microphone record permission, and the result is `:granted` only when both
are granted.
On Android 11 and later, if `MobSpeech.available?()` returns `false` even
though a recogniser is installed, add the package-visibility query to
`AndroidManifest.xml` (inside `<manifest>`):
```xml
<queries>
<intent><action android:name="android.speech.RecognitionService" /></intent>
</queries>
```
## Usage
```elixir
# Ask once (the platform engine needs [:speech]).
socket = Enum.reduce(MobSpeech.permissions(), socket, &Mob.Permissions.request(&2, &1))
# Hold-to-talk: listen on press, stop on release.
socket = MobSpeech.listen(socket, language: "en-US")
socket = MobSpeech.stop(socket) # deliver the final
socket = MobSpeech.cancel(socket) # abort, no final
def handle_info({:speech, :state, state}, socket), do: ... # :listening | :processing | :idle
def handle_info({:speech, :partial, text}, socket), do: ...
def handle_info({:speech, :final, text}, socket), do: ...
def handle_info({:speech, :error, reason}, socket), do: ...
```
| Function | Returns | Notes |
|---|---|---|
| `listen(socket, opts \\ [])` | socket | `language:` (BCP-47, default device locale), `prefer_offline:` (default `false`), `partial_results:` (default `true`), `stop_timeout_ms:` (default 2000 or the engine's), `engine:` (`:platform` or a module), `to:` (pid, default `self()`). Other options go to the engine. Listening again cancels the previous session. |
| `stop(socket)` | socket | `:processing`, then a final (or an error), then idle. No-op once the recognition has ended. |
| `cancel(socket)` | socket | idle, no final. |
| `available?(engine \\ :platform)` | boolean | `false` on a host build without the NIF. On iOS it stays `false` until speech recognition is authorised. |
| `permissions(engine \\ :platform)` | `[atom]` | capabilities to request first. |
A press shorter than ~300 ms has nothing to recognise. Android's recogniser
usually answers it with `:client` or `:no_speech`, so hold-to-talk apps should
cancel very short presses themselves. On Android, the platform engine also
takes `silence_ms:` (default 10 000), so a pause while the button is held
doesn't end the recognition.
These guarantees hold for every engine:
- After a final or an error you get exactly one `{:speech, :state, :idle}`,
and nothing else for that `listen`.
- `:processing` is sent when you call `stop/1`. It is never sent when the
recogniser endpoints by itself, because during hold-to-talk the button is
still held.
- The recogniser may finish by itself before you call `stop/1`. In that case
final and idle arrive first, and the later `stop/1` does nothing.
- If the final is empty, or the recogniser reports `:no_speech` after
partials, you get `{:speech, :final, last_partial}`.
- After `stop/1`, if the recogniser doesn't answer within `stop_timeout_ms`,
it is cancelled and you get the last partial (or `:no_speech`). The Google
recogniser can take 10–20 s to report after `stopListening`.
- If a permission is missing, `listen/2` still returns the socket. The screen
then gets `{:speech, :error, :permission}` followed by idle.
Error reasons: `:no_speech`, `:language`, `:permission`, `:service_permission`,
`:network`, `:audio`, `:busy`, `:client`, `:server`, `:unavailable`,
`:too_many_requests`, or `{:unknown, code}`. What each means:
- `:language`: no model or language pack for the locale.
- `:service_permission`: on Android, the recognition service (the Google app)
has no microphone access itself, even though your app does.
### Testing without a microphone
```elixir
MobSpeech.listen(socket,
engine: MobSpeech.Engine.Fake,
script: [:listening, {:partial, "hello"}, {:wait, 100}, {:partial, "hello world"}],
on_stop: [{:final, ""}] # empty final → you get "hello world"
)
```
## Development
```bash
mix setup # deps + git hooks (format, credo --strict, compile; tests when mix.exs changes)
mix test
```
## License
MIT