-
Notifications
You must be signed in to change notification settings - Fork 1.6k
en plugin dev components parser
The Parser component allows plugins to provide document parsing capabilities for LangBot. When a user uploads a document to a knowledge base, LangBot invokes the parser before the Knowledge Engine's ingest step, extracting structured text from binary files such as PDF, Word, Markdown, etc.
Relationship between Parser and KnowledgeEngine:
- Parser is responsible for converting files to text (file → text)
- KnowledgeEngine is responsible for indexing and retrieving text (text → chunks → vectors)
If the Knowledge Engine already has native document parsing capabilities (declared DOC_PARSING capability), users can choose to use the Knowledge Engine's built-in parsing or an external Parser plugin.
A single plugin can add any number of parsers. Execute the command lbp comp Parser in the plugin directory and follow the prompts to enter the parser configuration.
➜ MyParserPlugin > lbp comp Parser
Generating component Parser...
Parser name: pdf_parser
Parser description: A PDF document parser
Component Parser generated successfully.This will generate pdf_parser.yaml and pdf_parser.py files in the components/parser/ directory. The .yaml file defines the parser's basic information and supported MIME types, and the .py file is the handler for this parser:
➜ MyParserPlugin > tree
...
├── components
│ ├── __init__.py
│ └── parser
│ ├── __init__.py
│ ├── pdf_parser.py
│ └── pdf_parser.yaml
...apiVersion: v1 # Do not modify
kind: Parser # Do not modify
metadata:
name: pdf_parser # Parser name, used to identify this parser
label:
en_US: PDF Parser # Parser display name, shown in LangBot's UI, supports multiple languages
zh_Hans: PDF 解析器
description:
en_US: 'A PDF document parser'
zh_Hans: 'PDF 文档解析器'
spec:
supported_mime_types: # Declare supported file MIME types
- application/pdf
execution:
python:
path: pdf_parser.py # Parser handler, do not modify
attr: PdfParser # Class name of the parser handler, consistent with the class name in pdf_parser.pysupported_mime_types declares the file types this parser supports. Common MIME types:
| MIME Type | Description |
|---|---|
application/pdf |
PDF documents |
application/vnd.openxmlformats-officedocument.wordprocessingml.document |
Word documents (.docx) |
text/markdown |
Markdown files |
text/plain |
Plain text files |
text/html |
HTML files |
The following code will be generated by default (components/parser/<parser_name>.py). You need to implement the parse method.
from langbot_plugin.api.definition.components.parser.parser import Parser
from langbot_plugin.api.entities.builtin.rag.models import (
ParseContext,
ParseResult,
TextSection,
)
class PdfParser(Parser):
"""Parser component for extracting text from files."""
async def parse(self, context: ParseContext) -> ParseResult:
"""Parse a file and extract structured text.
Args:
context: Contains file_content (bytes), mime_type, filename, and metadata.
Returns:
ParseResult with extracted text and optional structured sections.
"""
# TODO: Implement parsing logic
text = context.file_content.decode("utf-8", errors="replace")
return ParseResult(
text=text,
sections=[
TextSection(
content=text,
heading=context.filename,
level=0,
),
],
metadata={
"filename": context.filename,
"mime_type": context.mime_type,
},
)The parse method is called when a document is uploaded to a knowledge base (before Knowledge Engine ingestion):
async def parse(self, context: ParseContext) -> ParseResult:ParseContext contains the following information:
class ParseContext(pydantic.BaseModel):
file_content: bytes # Raw file bytes (read by LangBot from storage)
mime_type: str # Detected MIME type of the file
filename: str # Original filename
metadata: dict[str, Any] # Extra metadata from FileObjectParseResult should return the parsing result:
class ParseResult(pydantic.BaseModel):
text: str # Full extracted plain text
sections: list[TextSection] = [] # Structured sections (optional)
metadata: dict[str, Any] = {} # Parsing metadata (e.g., page_count, language)TextSection represents a section of text extracted from the document:
class TextSection(pydantic.BaseModel):
content: str # Section text content
heading: str | None = None # Section heading
level: int = 0 # Nesting level
page: int | None = None # Source page number (for PDF, etc.)
metadata: dict[str, Any] = {} # Additional section metadataWhen a user uploads a document, LangBot determines the parsing flow as follows:
- If the user selects an external Parser plugin, LangBot first calls the Parser's
parsemethod, then passes the result to the Knowledge Engine'singestmethod viaIngestionContext.parsed_content. - If the Knowledge Engine declares
DOC_PARSINGcapability and the user does not select an external parser, the Knowledge Engine handles document parsing on its own.
KnowledgeEngine can check IngestionContext.parsed_content to determine whether pre-parsed content is available:
async def ingest(self, context: IngestionContext) -> IngestionResult:
if context.parsed_content:
# Use pre-parsed content from external Parser
text = context.parsed_content.text
sections = context.parsed_content.sections
else:
# Parse the document internally
file_bytes = await self.plugin.get_knowledge_file_stream(context.file_object.storage_path)
text = file_bytes.decode('utf-8')
...Before invoking a parser from another plugin, you can use self.plugin.list_parsers to discover the parsers currently available on the host:
parsers = await self.plugin.list_parsers(mime_type="application/pdf")
# Each item includes plugin_id, plugin_author, plugin_name, name, description, supported_mime_typesIf the list is empty, no connected Parser plugin currently supports that MIME type.
After obtaining plugin_author and plugin_name, you can call self.plugin.invoke_parser:
parser = parsers[0]
result = await self.plugin.invoke_parser(
plugin_author=parser["plugin_author"],
plugin_name=parser["plugin_name"],
storage_path=context.file_object.storage_path,
mime_type=context.file_object.metadata.mime_type,
filename=context.file_object.metadata.filename,
metadata={},
)
# result is a dict containing text, sections, metadataA single Parser is used by several knowledge bases (different knowledge bases of one installation upload files of different MIME types), and the platform storage key carries no knowledge-base dimension:
- Parse caches and temporary files must be isolated by knowledge-base ID, with the ID in the key; never cache them on the instance.
-
ParseContextand intermediate parse results are valid only for a single parse; never store them on instance fields or module-level variables. - A parse failure affects only that upload of the current knowledge base; it must not make parsing unavailable for the other knowledge bases.
- If you cache models, parsers or temporary directories per installation, release them in
on_installation_revoked(binding)and make sure temporary files can be cleaned up.
See Certified plugins and shared runtime for the full specification.
After creation, execute the command lbp run in the plugin directory to start debugging. Then in LangBot:
- Go to the "Knowledge Base" page
- Select a knowledge base and enter document management
- When uploading a file, select your plugin's parser in the parser selector
- After uploading, verify that the document is correctly ingested
Automatically synchronized from langbot-app/langbot-docs.
简体中文
指南
开发者
- 插件开发
- 插件 SDK API
- 核心开发
文章
- 浏览
- 产品动态
- 技术解析
- 教程与集成
- 公告
API 参考
- Service API
English
Guides
- Quick Start
- Installation
-
Configure Bots
- Bots
- Discord
- Telegram
- Slack
- Mattermost
- LINE
- Web Page Bot
- HTTP Bot
- KOOK
- Feishu
- DingTalk
- WeChat Official Account
- QQ (OneBot v11)
- Satori (QQ & Multi-Platform)
- QQ Official Bot
- WeCom (Enterprise WeChat)
- AI Configuration
- Advanced Operations
- Using Plugins
Developers
-
Plugin Development
- Plugin Development Tutorial
- Completing Plugin Configuration Information
- Plugin Directory Structure
- Component Development
- Code Style Guide
- Certified plugins and shared runtime
- Publish Plugin
- Migration Guide
- Plugin SDK API
- Core Development
Articles
- Browse
- Product Updates
- Engineering
-
Tutorials & Integrations
- LangTARS: Open-Source AI Agent for Remote PC Control — Works with Dify, n8n & 10+ Messaging Platforms
- How to Connect DeepSeek R1 to WeChat, Discord & Telegram in 5 Minutes (FREE)
- Deploy Your Own AI Bot to Discord, Telegram & WeChat in 5 Minutes
- Finally Got My Dify Agent Working in Discord, Telegram and Slack
- How I Built a Multi-Platform AI Bot with Langflow's Drag-and-Drop Workflows
- How I Built a Multi-Platform AI Chatbot with n8n and LangBot
- LangBot 4.6.0 External Knowledge Base Tutorial: Integrating Dify with LangBot for RAG-powered Conversations
- Announcements
API Reference
- Service API
Other pages
日本語
ガイド
開発者
- プラグイン開発
- プラグイン SDK API
- コア開発
記事
- 一覧
- 製品アップデート
- エンジニアリング
-
チュートリアルと連携
- LangTARS:Dify・n8n と連携するオープンソース PC 操作 Agent
- DeepSeek R1 を WeChat・Discord・Telegram に5分で接続する方法
- AI Bot を Discord・Telegram・WeChat に5分でデプロイ
- Dify Agent を Discord・Telegram・Slack で動かす
- Langflow のドラッグ&ドロップでマルチプラットフォーム AI Bot を構築
- n8n と LangBot でマルチプラットフォーム AI Chatbot を構築
- LangBot 4.6.0 外部ナレッジベース入門:Dify と連携した RAG 会話
- お知らせ
API リファレンス
- Service API