Every method below is accessed through self.capability_worker (the SDK) or self.worker (the Agent). This is the complete toolkit for building any Ability.
Twenty essential SDK methods
Bonus methods
delete_file() · get_audio_recording_length() · flush_audio_recording() · send_data_over_websocket() · send_devkit_action() · get_token() · get_region_data() · stream_init() / stream_end() · create_key() / update_key() / delete_key() / get_single_key() / get_all_keys() · update_personality_agent_prompt() · exec_local_command() · session_tasks.sleep() · session_tasks.get()
OpenRouter models
Use OpenRouter (openrouter.ai) as a single API endpoint to access any model. Pick by job: fast/cheap for routing, multimodal for audio, high-quality for user-facing responses.
Mix models in a single Ability. Use fast/cheap (Gemini Flash, Haiku, GPT-4o-mini) for intent routing and keyword extraction. Use quality models (Claude Sonnet, GPT-4o) for user-facing spoken responses. Use multimodal (Gemini Flash/Pro) for audio analysis.
Battle-tested prompt patterns
Each prompt below is designed for voice output — short, spoken, no markdown.
1. Intent router (JSON classification)
Use with text_to_text_response(). Always strip markdown fences before parsing JSON.
2. Persona system prompt (voice character)
Use as system_prompt parameter. Keep persona prompts specific about length, format, and forbidden phrases.
3. Audio analysis — Pass 1 (general)
Use with an OpenRouter audio-capable model (Gemini Flash/Pro). Send alongside base64 WAV.
4. Audio analysis — Pass 2 (specific with context)
Inject Pass 1 results as context. The two-pass pattern hides latency while providing deep answers.
5. Conversational response (with history)
Inject accumulated analysis + full chat history. Context compounds with every turn.
6. LLM-driven time parser (alarm pattern)
Loop up to 6 rounds. If response starts with QUESTION:, ask the user and continue.
Turns stream-of-consciousness rambling into structured, organized output.
8. Restart vs continue intent detection
Two-tier approach: check fast keywords first, fall back to LLM only for ambiguous input.
9. Contextual voice assistant
Inject user context (name, location, time) for natural, personalized responses.
10. Farewell / exit summary
Generate a contextual goodbye instead of a generic sign-off. Makes exits feel natural.
Architecture patterns
Ability categories
See Ability Types for the full breakdown.
File structure
main.py vs background.py
Core patterns
The loop template (multi-turn conversation)
Greet → loop (listen → process → respond) → exit on command. Most common pattern for interactive Abilities.
The two-pass analysis pattern
Pass 1 fires in background immediately (general analysis). While it runs, the Ability talks to the user. Pass 2 fires with Pass 1 context injected, answering the user’s specific question from depth.
- Pass 1: fire-and-forget via
session_tasks.create(asyncio.to_thread(run_general))
- Talk to user while Pass 1 runs (hides 10–15s of latency)
- Pass 2: inject Pass 1 results as context, answer the specific question
- Each follow-up turn fires a background re-analysis, enriching future turns
The rolling window pattern (ambient audio)
For always-on audio monitoring. Continuously record, slice the last N seconds, send to model on a fixed cadence. Fire-and-forget — never await inside the loop.
- 10-second window, 3-second refresh cadence
- API call fires as background task, poll loop never waits
- Responses arrive asynchronously and log themselves
The coordination pattern (main.py + background.py)
Main writes data to persistent file storage. Background polls that file on a timer and acts on it. This is how alarms, reminders, and scheduled tasks work.
main.py: parse user input, write to JSON file, resume_normal_flow()
background.py: poll file every 15–30 seconds, check conditions, act
- Use delete + write for JSON files — append corrupts JSON
- Call
send_interrupt_signal() before speaking from a daemon
The pending state pattern (multi-step collection)
Track what info you’re waiting for with a dictionary. Each loop iteration checks pending state first and routes input to the correct handler.
Sandbox rules
Breaking these rules will fail the Ability scanner.
- Never write
register_capability() by hand — always use the platform tag
- No
import os, no import json at the top level outside the register block
- No raw
open() — use play_from_audio_file() for audio, the file storage API for data
- No
signal module — even in docstrings or comments, the scanner catches it
- Always call
resume_normal_flow() on every exit path in main.py
- Use
session_tasks.sleep() and session_tasks.create() — not raw asyncio
- Wrap all blocking HTTP calls in
asyncio.to_thread()
- No
print() — use editor_logging_handler
- Blocked imports:
redis, connection_manager, user_config, exec(), eval(), pickle
Voice UX best practices
- Keep
speak() to 1–2 sentences. This is voice, not text
- Fill the silence: say “One sec” before any API call over 1 second
- Read your
speak() strings out loud before shipping
- Handle messy voice input: use the LLM to extract clean data from noisy transcription
- Offer exit at every loop iteration: check for “done”, “stop”, “quit”, etc.
- Use
run_confirmation_loop() before destructive actions (send, delete, cancel)
- Idle detection: 1 empty response = keep going, 2 in a row = offer to leave
- Namespace your filenames:
smarthub_prefs.json not data.json
- JSON persistence: always delete + write (append corrupts JSON)
- API calls: always set
timeout=10, wrap in try/except, speak errors to user