Convert scientific documents into structured, searchable markdown using AI-powered vision analysis.
Built for pharmaceutical research — handles conference posters, patent filings, presentation slides, and corporate presentations (PPTX + PDF) out of the box.
| Pipeline | Input | Output |
|---|---|---|
| Poster | Conference poster PDFs | Structured markdown with figures, sections, and metadata |
| Patent | Patent filing PDFs (WO/EP/US) | Claims, chemical structures, SMILES, executive summaries |
| Talk | Slide-based presentation PDFs | Slide-by-slide extraction with narrative summaries |
| Presentation | PPTX + PDF presentations | Native text extraction, per-slide image descriptions, action items, chemical structures |
| Paper | Scientific research paper PDFs | Full-text extraction with sections, figures, and summary |
| CI | Competitive intelligence PDFs + PPTXs | Company/asset structured markdown, deal & regulatory event extraction |
| Outlook | Outlook email folders | Threaded conversations, processed attachments, persistent contact book |
Each pipeline produces a self-contained .md file with YAML frontmatter, making the output easy to search, filter, and integrate into knowledge bases.
First time? See the full setup guide for step-by-step installation from scratch (Miniconda, packages, credentials, troubleshooting).
conda create -n ds_env python=3.11 -y
conda activate ds_env
conda install -c conda-forge pdfplumber pymupdf pillow -y
conda install pandas openpyxl -y
conda install -c conda-forge tesseract pytesseract -y
pip install anthropic[Environment]::SetEnvironmentVariable("ANTHROPIC_AUTH_TOKEN", "your-token", "User")
[Environment]::SetEnvironmentVariable("ANTHROPIC_BASE_URL", "your-proxy-url", "User")Restart your terminal after setting these.
Posters:
# with metadata
python poster_pipeline.py --input "path/to/poster_pdfs" --metadata "abstracts.xlsx"
# without metadata
python poster_pipeline.py --input "path/to/poster_pdfs"
# standardized filenames
python poster_pipeline.py --input "path/to/poster_pdfs" --naming standardized --conference AACR --year 2026Patents:
python patent_pipeline.py --input "path/to/patent_pdfs"
# or a single file:
python patent_pipeline.py --single "path/to/WO2024123456.pdf"
# recursive scan with detailed naming:
python patent_pipeline.py --input "path/to/patent_pdfs" --recursive --naming detailedTalks:
python talk_pipeline.py --input "path/to/talk_pdfs" --metadata "abstracts.xlsx"
# or a single file:
python talk_pipeline.py --single "path/to/talk.pdf"Presentations:
python presentation_pipeline.py --input "path/to/presentations"
# text-only (no API calls):
python presentation_pipeline.py --input "path/to/presentations" --no-vision
# single file with dated naming:
python presentation_pipeline.py --single "path/to/file.pptx" --naming datedPapers:
python paper_pipeline.py --input "path/to/papers"
# single file:
python paper_pipeline.py --single "path/to/paper.pdf"Competitive Intelligence:
python ci_pipeline.py --input "path/to/ci_docs"
# custom output location:
python ci_pipeline.py --input "path/to/ci_docs" --output "output_ci"Outlook Emails:
python outlook_pipeline.py --folder "Inbox/CI Reports"
# limit to recent emails, skip attachments:
python outlook_pipeline.py --folder "Inbox" --limit 50 --no-attachments
# reprocess everything:
python outlook_pipeline.py --folder "Inbox/Projects" --no-skipAll pipelines share a consistent CLI interface:
| Flag | Description |
|---|---|
--input |
Input folder containing documents (aliases: --sharepoint, --talks) |
--output |
Output directory for markdown files |
--single |
Process a single file instead of a folder |
--recursive |
Recursively search subfolders for files |
--no-skip |
Reprocess files that already exist (default: skip existing) |
--naming |
Output filename scheme (options vary per pipeline) |
--verbose |
Enable debug logging |
Each pipeline also has specialized flags — see the pipeline-specific documentation for full options:
- Poster Pipeline flags —
--metadata,--force-ocr,--conference,--col-* - Patent Pipeline flags —
--no-vision,--claims-only,--budget,--ocr-engine - Talk Pipeline flags —
--metadata - Presentation Pipeline flags —
--no-vision - Outlook Pipeline flags —
--folder,--no-attachments,--limit
Each pipeline generates markdown files with:
- YAML frontmatter — structured metadata (dates, authors, scores, classifications)
- Executive summary — AI-generated overview of the document
- Full content — sections, claims, or slides extracted from the PDF
- Figures & visuals — descriptions of charts, structures, and diagrams
- Quality score — automated 0–10 quality assessment
Example output location:
output/
├── poster_1234.md
├── poster_1235.md
└── quality_log.txt
output_patents/
├── patent_WO2024123456A1.md
└── quality_log.txt
output_talks/
├── talk_04_ED03_Bunne_toward_virtual_patients.md
└── processing_log.txt
output_presentations/
├── presentation_ru_onc_operations_update_darmstadt.md
└── presentation_caris_discovery_non_con_apr.md
output_outlook/
├── processed_state.json
├── processing_log.tsv
├── people/ # global contact book
│ ├── carsten_schweer.yml
│ └── alice_smith.yml
└── Inbox/ # mirrors Outlook folder hierarchy
└── CI Reports/
├── thread_project_alpha_update/
│ ├── thread.md
│ └── attachment_report.md
└── thread_meeting_notes_q2/
└── thread.md
| Poster | Patent | Talk | Presentation | Outlook | |
|---|---|---|---|---|---|
| Input | Single-page poster PDF | Multi-page patent PDF (50–300+ pages) | Multi-slide presentation PDF (screenshots) | PPTX or PDF slide decks | Outlook email folder |
| Output | Sections (Methods, Results, Conclusions) | Patent sections + claims + chemical data | Slide-by-slide content + narrative summary | Slide content + action items + metrics | Threaded conversations + processed attachments |
| Text extraction | Native PDF + OCR + Vision AI | Native PDF + Tesseract OCR (scanned) + selective Vision AI for garbled pages | OCR + Vision AI only (no extractable text) | Native PPTX (python-pptx) or PyMuPDF + pdfplumber tables + Vision AI fallback | Outlook COM (plain text / HTML conversion) |
| Metadata source | Excel spreadsheet (optional) | Extracted from the PDF itself | Excel spreadsheet (optional) | Extracted from the file itself | Extracted from email headers |
| Quality gate | Yes — skips FAIR/POOR | No — all patents saved | No — all talks saved | No — all presentations saved | No — all threads saved |
| Poster | Patent | Talk | Presentation | Outlook | |
|---|---|---|---|---|---|
| Processing time | 1.5–4 min | 2–3 min (text+vision), 3–5 min (scanned+Tesseract), ~15s text-only | 1.5–4 min | <1s text-only, 30–90s with vision, 2–4 min image-heavy | ~1s per email (text only), +pipeline time for attachments |
| API calls per doc | 4 + N figures | 10–48 (text-native), 8–35 (scanned+Tesseract) | 4–9 (scales with slide count) | 1–3 (smart gating) + 5–15 (image enrichment), 0 text-only | 0 (email body), varies for attachments |
| Token usage per doc | ~30K–60K | ~40K–120K | ~18K–35K | ~8K–20K (text-only), ~50K–120K (image-heavy) | 0 (email body), varies for attachments |
| Vision AI pages | All pages (mandatory) | ~10% of pages (selective) | All slides (mandatory) | Smart gating: global analysis + per-slide image enrichment (requires PowerPoint for PPTX) | None (attachments only) |
| Text-only mode | No | Yes (--no-vision) |
No | Yes (--no-vision) |
Yes (default for email body) |
| Claims-only mode | No | Yes (--claims-only, ~5s) |
No | No | No |
| Concurrency | Up to 5 figure workers | Up to 3 batch workers | Up to 5 batch workers | Sequential | Sequential |
| Capability | Poster | Patent | Talk | Presentation | Outlook |
|---|---|---|---|---|---|
| Two-stage figure analysis | x | ||||
| Chemical structure extraction (SMILES) | x | x | |||
| Claims dependency tree | x | ||||
| Semantic classification (target, mechanism, modality) | x | ||||
| Hybrid OCR (Tesseract + Vision AI) | x | ||||
| Text quality scoring & auto-repair | x | ||||
| OCR pre-pass as RAG context | x | x | |||
| Abstract matching from metadata | x | x | |||
| Native PPTX text extraction | x | ||||
| Per-slide image enrichment (auto-describes figures/screenshots) | x | ||||
| Smart Vision AI gating (skip when not needed) | x | ||||
| Language detection (EN/DE) + English output | x | ||||
| Classification detection (3-signal) | x | ||||
| Action items & metrics extraction | x | ||||
| Conditional summary (content-type aware) | x | ||||
| Conversation threading (ConversationID + subject) | x | ||||
| Incremental processing (EntryID state tracking) | x | ||||
| Attachment routing to sub-pipelines | x | ||||
| HTML-to-markdown email body conversion | x | ||||
| Standardized naming schemes | x | x | x | ||
| Executive summary | x | x | x | x | |
| Quality scoring | x | x | x | x |
- Poster — single-page conference posters with figures, methods, and results sections
- Patent — multi-page patent filings (WIPO, EPO, USPTO) with claims, chemical structures, and experimental data
- Talk — slide-based presentations captured as PDF screenshots (no extractable text)
- Presentation — PPTX or PDF slide decks with native text (corporate meetings, scientific presentations, agendas); supports English and German content
- Outlook — email conversations from any Outlook folder; automatically processes PDF/PPTX/DOCX/XLSX attachments through the appropriate pipeline
- Text Extraction — native text via PyMuPDF/pdfplumber/python-pptx, with OCR fallback (Tesseract) for posters/talks/scanned patents
- Page Rendering — pages rendered to images using PyMuPDF (PDFs) or PowerPoint COM (PPTX) when Vision AI is needed
- Vision AI — Claude analyzes page images for figures, chemical structures, and garbled text (smart gating skips this when not needed)
- Structuring — extracted content is parsed into logical sections
- Summarization — AI generates executive summaries and key findings (always in English)
- Quality Scoring — automated scoring flags documents that may need manual review
| Guide | Description |
|---|---|
| Setup Guide | Full installation from scratch — Miniconda, packages, credentials, troubleshooting |
| Poster Pipeline | Deep dive — architecture, processing stages, figure analysis, quality scoring |
| Patent Pipeline | Deep dive — claims parsing, SMILES extraction, text quality repair, semantic classification |
| Talk Pipeline | Deep dive — slide extraction, OCR pre-pass, summary generation, batch processing |
| Presentation Pipeline | Deep dive — PPTX/PDF extraction, classification detection, action items, naming schemes |
| Outlook Pipeline | Deep dive — email threading, incremental processing, attachment routing |
- Python 3.11+ (via Conda)
- Tesseract OCR (poster/talk/patent pipelines — required for poster/talk, optional for patent scanned PDFs)
- python-pptx, PyMuPDF, pdfplumber (presentation/paper pipelines)
- python-docx (outlook pipeline — DOCX attachment extraction)
- pyyaml (outlook pipeline — contact book YAML read/write; usually pre-installed with conda)
- Microsoft PowerPoint (optional — enables Vision AI for PPTX slide rendering)
- Microsoft Outlook (outlook pipeline — must be running; no OAuth/admin required)
- Anthropic API access (Claude Sonnet 4.6)
- Windows (uses Windows Registry for credential loading; Outlook COM requires Windows)
Note:
paper_pipeline.pyandci_pipeline.pydo not yet have dedicated README files — see inline--helpfor their CLI options.