A Telegram message search bot with full-text search and Chinese word segmentation support.
Note: This is a Rust rewrite of the original Python implementation. The Python version is deprecated and no longer maintained.
Telegram's built-in search functionality is limited, especially for CJK languages like Chinese where proper word segmentation is crucial. TG Searcher provides a comprehensive solution by indexing your Telegram messages locally and offering powerful full-text search capabilities through a bot interface.
- Full-text search powered by Tantivy with Chinese word segmentation (jieba)
- Real-time message indexing as new messages arrive
- Support for message edits and deletions
- Advanced search syntax (AND, OR, NOT, wildcards, phrases)
- Chat-specific search filtering
- Pagination through search results
- Admin commands for index management
- In-memory state management (no Redis dependency)
- High performance async implementation in Rust
main.rs (Orchestrator)
├─> Config (YAML parsing, validation)
├─> ClientSession (Telegram user login, chat cache)
├─> BackendBot (message monitoring, indexing)
│ └─> Indexer (Tantivy search with jieba)
└─> BotFrontend (bot commands, user interaction)
└─> Storage (pagination state)
- ClientSession: Manages Telegram user sessions with chat name caching
- BackendBot: Monitors chats and indexes messages in real-time
- Indexer: Tantivy-based search engine with Chinese tokenization
- BotFrontend: Telegram bot for user commands and search queries
- Storage: In-memory state management for pagination
- Rust 1.70 or higher (2021 edition)
- Telegram API credentials from https://my.telegram.org
- A Telegram bot token from @BotFather
cargo build --releaseThe binary will be available at target/release/tg-searcher.
- Copy the example configuration:
cp searcher.yaml.example searcher.yaml- Edit
searcher.yamlwith your credentials:
common:
api_id: 12345678
api_hash: "your_api_hash_here"
runtime_dir: "./tg_searcher_data"
# proxy: "socks5://localhost:1080" # Optional
sessions:
- name: "my_session"
phone: "+1234567890"
backends:
- id: "backend1"
use_session: "my_session"
config:
monitor_all: false
excluded_chats: []
frontends:
- id: "frontend1"
use_backend: "backend1"
config:
bot_token: "1234567890:ABCdefGHIjklMNOpqrsTUVwxyz"
admin_id: 123456789
page_len: 10
private_mode: false
private_whitelist: []api_id: Telegram API IDapi_hash: Telegram API hashruntime_dir: Directory for storing sessions and indexesproxy: Optional SOCKS5 proxy (e.g.,socks5://localhost:1080)
name: Unique session identifierphone: Phone number for authentication
id: Unique backend identifieruse_session: Reference to session namemonitor_all: Monitor all chats (default: false)excluded_chats: Chat IDs to exclude when monitor_all is true
id: Unique frontend identifieruse_backend: Reference to backend IDbot_token: Bot token from @BotFatheradmin_id: Telegram user ID of the adminpage_len: Results per page (default: 10)private_mode: Restrict access to whitelisted usersprivate_whitelist: List of allowed user IDs
# Normal mode
./target/release/tg-searcher -c searcher.yaml
# Clear existing index and start fresh
./target/release/tg-searcher --clear -c searcher.yaml
# Enable debug logging
./target/release/tg-searcher --debug -c searcher.yamlOn first run:
- Authenticate the user session (enter verification code sent to Telegram)
- If 2FA is enabled, enter your password
- Send
/startto your bot to initialize it
/search <query>- Search messages (or just type your query directly)/random- Get a random indexed message/chats [keyword]- List monitored chats, optionally filtered by keyword
/stat- Show index statistics/download_chat [min=N] [max=N] [CHAT...]- Download and index chat history/monitor_chat [CHAT...]- Add chat to monitoring list/clear [all|CHAT...]- Clear index (all or specific chats)/find_chat_id <keyword>- Find chat IDs by name/refresh_chat_names- Refresh chat name cache
The search supports Whoosh query syntax:
"foo bar"- Search for exact phrasefoo AND bar- Both terms must appearfoo OR bar- Either term can appearNOT foo- Exclude messages containing "foo"foo*- Wildcard (any characters after "foo")foo?- Single character wildcard
Examples:
"hello world"- Messages containing exact phrase "hello world"hello AND world- Messages containing both "hello" and "world"NOT spam AND (buy OR sell)- Messages without "spam" but with "buy" or "sell"
Use /chats to list available chats. Click a chat button, then reply to that message with your search query to search only within that chat.
src/
├── main.rs # Application entry point and orchestration
├── config.rs # YAML configuration parsing and validation
├── types.rs # Core types and error definitions
├── utils.rs # Utility functions (escape, share_id, etc.)
├── storage.rs # Storage trait and in-memory implementation
├── indexer.rs # Tantivy indexer with Chinese tokenization
├── session.rs # Telegram session management (TODO: grammers)
├── backend.rs # Message indexing backend (TODO: event handlers)
└── frontend.rs # Bot command handlers (TODO: bot integration)
# Run all tests
cargo test
# Run with output
cargo test -- --nocapture
# Run specific test
cargo test test_indexer_basic_operationsCurrent test coverage: 12 passing tests covering:
- Configuration parsing and validation
- Proxy parsing
- Utility functions
- In-memory storage
- Tantivy indexer operations
- CLI argument parsing
# Check compilation
cargo check
# Run linter
cargo clippy
# Format code
cargo fmtEngine: Tantivy 0.25
Tokenizer: jieba-rs for Chinese word segmentation
Schema:
content: Full-text indexed with Chinese analyzerurl: Unique identifier (e.g.,https://t.me/c/1234567890/123)chat_id: Indexed for filteringpost_time: Indexed and fast-sorted (descending)sender: Stored for display
- Indexing: Approximately 50MB heap for writer
- Search: Sub-second queries for millions of messages
- Memory: Lock-free concurrent access using DashMap
- Storage: On-disk persistent Tantivy index
tokio: Async runtimegrammers-client: Telegram client (integration in progress)tantivy: Full-text search enginejieba-rs: Chinese text segmentationserde: Configuration serializationtracing: Structured loggingclap: CLI argument parsingdashmap: Concurrent hashmapanyhow/thiserror: Error handling
- Core infrastructure (types, config, utils, storage)
- Tantivy indexer with Chinese tokenization
- Session management structure
- Backend bot with event handler structure
- Frontend bot with all command handlers
- Main orchestration and initialization
- Comprehensive test suite
- Documentation
- grammers Telegram client integration
- Session connection and authentication
- Event handlers (NewMessage, MessageEdited, MessageDeleted)
- Bot client and command registration
- Message sending and callback handling
- End-to-end integration testing
- Docker deployment setup
- Performance benchmarks
- Migration tool from Python version
- Performance: Native Rust implementation with async I/O
- Type Safety: Compile-time correctness guarantees
- Memory Efficiency: No GIL, superior concurrency model
- Search Engine: Tantivy (faster than Whoosh)
- Architecture: Cleaner separation of concerns
- Dependencies: Fewer runtime dependencies
- Redis: Replaced with in-memory storage using trait-based design for future extensibility
The bot commands and user experience remain identical to the Python version for easy migration.
./tg-searcher --clear -c searcher.yaml# Remove session file and re-authenticate
rm tg_searcher_data/sessions/my_session.session
./tg-searcher -c searcher.yaml# Via command line
./tg-searcher --debug -c searcher.yaml
# Via environment variable
RUST_LOG=debug ./tg-searcher -c searcher.yamlContributions are welcome! Please:
- Fork the repository
- Create a feature branch
- Add tests for new functionality
- Ensure all tests pass (
cargo test) - Run code formatter (
cargo fmt) - Run linter (
cargo clippy) - Submit a pull request