A practical authentication workflow that can combine:
Face verification (InsightFace embeddings)
Speaker verification (SpeechBrain speaker embeddings)
Optional offline speech-to-text (STT) for capturing a name using Vosk
This repository is published in a privacy-safe way: the project structure is complete, but any sensitive datasets/logs/identities were removed or replaced with placeholders.
We typically run the complete system using:
python run_system.py --stt_name --vosk_model vosk-model-small-en-us-0.15 --name_seconds 6 --cooldown 10 --pc_ui_lang en--stt_nameEnables offline speech-to-text for capturing/confirming a spoken name.--vosk_model <folder>Path (or folder name) of the Vosk model directory.--name_seconds 6How many seconds to listen for the name.--cooldown 10Cooldown time between attempts.--pc_ui_lang enUI language.
- Multi-factor verification (Face + Voice)
- Offline STT option (Vosk)
- Privacy-safe dataset templates
- UI prompts (AR/EN)
- Modular scripts
This repo intentionally excludes any private data.
- dataset/ exists as structure only
- db files are placeholders
- logs are empty
.
├─ run_system.py # Main runner (full system)
├─ main.py # FastAPI server (if used by your flow)
├─ pc_client.py # PC-side logic (STT + UI prompts)
├─ verify_fusion.py # Fusion logic (face + voice)
├─ face_model_insightface.py # Face embeddings using InsightFace
├─ voice_model.py # Speaker embeddings using SpeechBrain
├─ db/
│ ├─ teachers.json # Template (placeholder)
│ └─ pending.json # Template (placeholder)
├─ dataset/ # Template only (no private media)
├─ logs/
│ └─ attempts.jsonl # Empty placeholder
└─ assets/prompts/ # Audio/text prompts (AR/EN)
- Python 3.9+
- numpy, requests, tqdm
- opencv-python
- torch, torchaudio
- speechbrain
- insightface
- fastapi, uvicorn
- optional: vosk
python -m venv .venv
.venv\Scripts\activatepip install -U pip
pip install numpy requests tqdm opencv-python
pip install torch torchaudio
pip install speechbrain insightface
pip install fastapi uvicorn
pip install voskA folder like vosk-model-small-en-us-0.15 is an external pre-trained model downloaded from a third-party source. To avoid licensing and ownership issues, and to reduce repo size, the model is not committed here.
Download any compatible Vosk model, whether English or another language.
Place the model folder inside the project directory, for example:
Authentication-System/
└─ vosk-model-small-en-us-0.15/
No problem. Just pass its name or path using --vosk_model:
python run_system.py --stt_name --vosk_model vosk-model-small-en-us-0.22 --name_seconds 6 --cooldown 10 --pc_ui_lang enAlternatively, you can change the default in the code where the argument is defined by searching for --vosk_model in run_system.py.
From the project root:
python run_system.py --stt_name --vosk_model vosk-model-small-en-us-0.15 --name_seconds 6 --cooldown 10 --pc_ui_lang enRun without STT, if supported by your setup:
python run_system.py --vosk_model vosk-model-small-en-us-0.15 --cooldown 10 --pc_ui_lang enSwitch UI language, if you have prompts for Arabic:
python run_system.py --stt_name --vosk_model vosk-model-small-en-us-0.15 --name_seconds 6 --cooldown 10 --pc_ui_lang arTo actually enroll or verify identities, you must provide your own data:
- Add your own images and audio into the expected
dataset/structure. - Fill your local teacher or user list, such as CSV or JSON templates, with your own IDs.
- Generate embeddings and databases according to the scripts used in your workflow.
- The system can’t find the Vosk model: Make sure the folder exists and the path matches the
--vosk_modelvalue. - Empty dataset or missing identities: This repo ships without private media by design. Add your own data locally.
- Repeated triggers or too many attempts: Increase
--cooldownto reduce back-to-back attempts.
This repository contains project code and placeholder templates only. Third-party models, such as Vosk, are governed by their original licenses and must be obtained separately.
If you build on this project, feel free to open an issue or submit a pull request.