Your voice, amplified
Chorus is an AI-powered Augmentative and Alternative Communication (AAC) tool designed to close the "Rate Gap" for non-verbal users. By shifting from reactive typing to active prediction, Chorus restores agency, speed, and vocal identity to those who need it most.
- Spoken Speech: ~150 words per minute.
- Traditional AAC: ~15 words per minute.
- Result: A fundamental disconnect that forces users into a passive, reactive role in conversations.
Chorus is not just a keyboard; it is an active listener. It uses a Multi-Agent Engine to fuse 7 real-time context signals, allowing it to predict what the user wants to say before they even touch the screen.
- The Ears (Listening): Uses
OpenAI Whisperto actively listen to the conversation partner and understand context. - The Scheduler (Time): Syncs with the user's daily itinerary to enable time-aware predictions (e.g., suggesting "Lunch" at noon).
- The Memory (History): Powered by Backboard.io, Chorus has "Object Permanence." It remembers names, past events, and friends.
- The Frequency (Habits): Uses
MongoDB Atlasto track selection habits, ranking a user's favorite words higher to save keystrokes. - The Filter (Context): Filters suggestions based on topic context (e.g., if talking about "Food" and the user types "P", it suggests "Pizza" but filters out "Paper").
- The Grammar (Syntax): Predicts the next part of speech (Noun vs Verb) to construct valid sentences.
- The Tone (Identity): Powered by ElevenLabs, Chorus injects emotional prosody (Joy, Sadness, Affectionate) into the synthetic voice.
- Frontend: Next.js, TypeScript, TailwindCSS
- Orchestrator: Google Gemini
- Voice Synthesis: ElevenLabs API
- Transcription: OpenAI Whisper
- Memory Vector DB: Backboard.io
- Database: MongoDB Atlas
- Image Generation: DALL-E 3 (for the "Infinite Icon" feature)
- Text Mode: For literate users. Features "Smart Type" providing sentence completions.
- Pictorial Mode: Mainly for non-verbal autistic kids. Features Infinite Icon (DALL-E 3) to generate custom symbols on the fly.
- Spark Mode: The "Anti-Passenger" tool. Suggests conversation starters to help users initiate dialogue.
Users can select an emotional intent (e.g., "Excited"). Chorus doesn't just read the text; it modifies pitch, cadence, and stability to make the voice sound excited.
To combat API latency, we implemented a semantic caching layer. If the conversation topic hasn't changed, Chorus serves instant, cached predictions instead of re-querying the LLM.
We integrate Backboard.io to provide the AI with object permanence. Chorus stores semantic vector embeddings of conversation history, allowing it to recall specific details—like a pet's name or facts about them from past conversations—ensuring the AI remembers the user's life history without needing reminders.
-
Clone the repo
git clone https://github.com/yourusername/chorus.git cd chorus -
Install dependencies
npm install
-
Set up Environment Variables Create a
.env.localfile with the following keys:ELEVENLABS_API_KEY= GEMINI_API_KEY= OPENAI_API_KEY= BACKBOARD_API_KEY= MONGO_DB= BACKBOARD_ASSISTANT_ID=
-
Run the development server
npm run dev
- 90% Keystroke Reduction: Validated that efficient prediction can turn a 15-tap sentence into a 2-tap confirmation.
- Infinite Icon: Successfully integrated DALL-E 3 to allow users to create new vocabulary items instantly.
- Emotional Resonance: Proved that AI voice synthesis can convey sadness, affection, and joy, effectively bridging the "Intellectual Gap."
- Offline Mode: Distilling the reasoning engine into local models (Gemini Nano) for non-internet use.
- Eye-Tracking Support: Optimizing the UI layout for gaze-based interaction.