Skip to content

About

AI-powered scientific document processing pipelines (patents, posters, talks)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

41 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Doc2MD

DocumentToMarkdown

Convert scientific documents into structured, searchable markdown using AI-powered vision analysis.

Built for pharmaceutical research — handles conference posters, patent filings, presentation slides, and corporate presentations (PPTX + PDF) out of the box.

What It Does

Pipeline Input Output
Poster Conference poster PDFs Structured markdown with figures, sections, and metadata
Patent Patent filing PDFs (WO/EP/US) Claims, chemical structures, SMILES, executive summaries
Talk Slide-based presentation PDFs Slide-by-slide extraction with narrative summaries
Presentation PPTX + PDF presentations Native text extraction, per-slide image descriptions, action items, chemical structures
Paper Scientific research paper PDFs Full-text extraction with sections, figures, and summary
CI Competitive intelligence PDFs + PPTXs Company/asset structured markdown, deal & regulatory event extraction
Outlook Outlook email folders Threaded conversations, processed attachments, persistent contact book

Each pipeline produces a self-contained .md file with YAML frontmatter, making the output easy to search, filter, and integrate into knowledge bases.

Quick Start

First time? See the full setup guide for step-by-step installation from scratch (Miniconda, packages, credentials, troubleshooting).

1. Setup Environment

conda create -n ds_env python=3.11 -y
conda activate ds_env

conda install -c conda-forge pdfplumber pymupdf pillow -y
conda install pandas openpyxl -y
conda install -c conda-forge tesseract pytesseract -y
pip install anthropic

2. Set Credentials

[Environment]::SetEnvironmentVariable("ANTHROPIC_AUTH_TOKEN", "your-token", "User")
[Environment]::SetEnvironmentVariable("ANTHROPIC_BASE_URL", "your-proxy-url", "User")

Restart your terminal after setting these.

3. Run a Pipeline

Posters:

# with metadata
python poster_pipeline.py --input "path/to/poster_pdfs" --metadata "abstracts.xlsx"
# without metadata
python poster_pipeline.py --input "path/to/poster_pdfs"
# standardized filenames
python poster_pipeline.py --input "path/to/poster_pdfs" --naming standardized --conference AACR --year 2026

Patents:

python patent_pipeline.py --input "path/to/patent_pdfs"
# or a single file:
python patent_pipeline.py --single "path/to/WO2024123456.pdf"
# recursive scan with detailed naming:
python patent_pipeline.py --input "path/to/patent_pdfs" --recursive --naming detailed

Talks:

python talk_pipeline.py --input "path/to/talk_pdfs" --metadata "abstracts.xlsx"
# or a single file:
python talk_pipeline.py --single "path/to/talk.pdf"

Presentations:

python presentation_pipeline.py --input "path/to/presentations"
# text-only (no API calls):
python presentation_pipeline.py --input "path/to/presentations" --no-vision
# single file with dated naming:
python presentation_pipeline.py --single "path/to/file.pptx" --naming dated

Papers:

python paper_pipeline.py --input "path/to/papers"
# single file:
python paper_pipeline.py --single "path/to/paper.pdf"

Competitive Intelligence:

python ci_pipeline.py --input "path/to/ci_docs"
# custom output location:
python ci_pipeline.py --input "path/to/ci_docs" --output "output_ci"

Outlook Emails:

python outlook_pipeline.py --folder "Inbox/CI Reports"
# limit to recent emails, skip attachments:
python outlook_pipeline.py --folder "Inbox" --limit 50 --no-attachments
# reprocess everything:
python outlook_pipeline.py --folder "Inbox/Projects" --no-skip

Common Flags

All pipelines share a consistent CLI interface:

Flag Description
--input Input folder containing documents (aliases: --sharepoint, --talks)
--output Output directory for markdown files
--single Process a single file instead of a folder
--recursive Recursively search subfolders for files
--no-skip Reprocess files that already exist (default: skip existing)
--naming Output filename scheme (options vary per pipeline)
--verbose Enable debug logging

Each pipeline also has specialized flags — see the pipeline-specific documentation for full options:

Output Structure

Each pipeline generates markdown files with:

  • YAML frontmatter — structured metadata (dates, authors, scores, classifications)
  • Executive summary — AI-generated overview of the document
  • Full content — sections, claims, or slides extracted from the PDF
  • Figures & visuals — descriptions of charts, structures, and diagrams
  • Quality score — automated 0–10 quality assessment

Example output location:

output/
├── poster_1234.md
├── poster_1235.md
└── quality_log.txt

output_patents/
├── patent_WO2024123456A1.md
└── quality_log.txt

output_talks/
├── talk_04_ED03_Bunne_toward_virtual_patients.md
└── processing_log.txt

output_presentations/
├── presentation_ru_onc_operations_update_darmstadt.md
└── presentation_caris_discovery_non_con_apr.md

output_outlook/
├── processed_state.json
├── processing_log.tsv
├── people/                         # global contact book
│   ├── carsten_schweer.yml
│   └── alice_smith.yml
└── Inbox/                          # mirrors Outlook folder hierarchy
    └── CI Reports/
        ├── thread_project_alpha_update/
        │   ├── thread.md
        │   └── attachment_report.md
        └── thread_meeting_notes_q2/
            └── thread.md

Pipeline Comparison

At a Glance

Poster Patent Talk Presentation Outlook
Input Single-page poster PDF Multi-page patent PDF (50–300+ pages) Multi-slide presentation PDF (screenshots) PPTX or PDF slide decks Outlook email folder
Output Sections (Methods, Results, Conclusions) Patent sections + claims + chemical data Slide-by-slide content + narrative summary Slide content + action items + metrics Threaded conversations + processed attachments
Text extraction Native PDF + OCR + Vision AI Native PDF + Tesseract OCR (scanned) + selective Vision AI for garbled pages OCR + Vision AI only (no extractable text) Native PPTX (python-pptx) or PyMuPDF + pdfplumber tables + Vision AI fallback Outlook COM (plain text / HTML conversion)
Metadata source Excel spreadsheet (optional) Extracted from the PDF itself Excel spreadsheet (optional) Extracted from the file itself Extracted from email headers
Quality gate Yes — skips FAIR/POOR No — all patents saved No — all talks saved No — all presentations saved No — all threads saved

Performance & Cost

Poster Patent Talk Presentation Outlook
Processing time 1.5–4 min 2–3 min (text+vision), 3–5 min (scanned+Tesseract), ~15s text-only 1.5–4 min <1s text-only, 30–90s with vision, 2–4 min image-heavy ~1s per email (text only), +pipeline time for attachments
API calls per doc 4 + N figures 10–48 (text-native), 8–35 (scanned+Tesseract) 4–9 (scales with slide count) 1–3 (smart gating) + 5–15 (image enrichment), 0 text-only 0 (email body), varies for attachments
Token usage per doc ~30K–60K ~40K–120K ~18K–35K ~8K–20K (text-only), ~50K–120K (image-heavy) 0 (email body), varies for attachments
Vision AI pages All pages (mandatory) ~10% of pages (selective) All slides (mandatory) Smart gating: global analysis + per-slide image enrichment (requires PowerPoint for PPTX) None (attachments only)
Text-only mode No Yes (--no-vision) No Yes (--no-vision) Yes (default for email body)
Claims-only mode No Yes (--claims-only, ~5s) No No No
Concurrency Up to 5 figure workers Up to 3 batch workers Up to 5 batch workers Sequential Sequential

Unique Capabilities

Capability Poster Patent Talk Presentation Outlook
Two-stage figure analysis x
Chemical structure extraction (SMILES) x x
Claims dependency tree x
Semantic classification (target, mechanism, modality) x
Hybrid OCR (Tesseract + Vision AI) x
Text quality scoring & auto-repair x
OCR pre-pass as RAG context x x
Abstract matching from metadata x x
Native PPTX text extraction x
Per-slide image enrichment (auto-describes figures/screenshots) x
Smart Vision AI gating (skip when not needed) x
Language detection (EN/DE) + English output x
Classification detection (3-signal) x
Action items & metrics extraction x
Conditional summary (content-type aware) x
Conversation threading (ConversationID + subject) x
Incremental processing (EntryID state tracking) x
Attachment routing to sub-pipelines x
HTML-to-markdown email body conversion x
Standardized naming schemes x x x
Executive summary x x x x
Quality scoring x x x x

When to Use Which

  • Poster — single-page conference posters with figures, methods, and results sections
  • Patent — multi-page patent filings (WIPO, EPO, USPTO) with claims, chemical structures, and experimental data
  • Talk — slide-based presentations captured as PDF screenshots (no extractable text)
  • Presentation — PPTX or PDF slide decks with native text (corporate meetings, scientific presentations, agendas); supports English and German content
  • Outlook — email conversations from any Outlook folder; automatically processes PDF/PPTX/DOCX/XLSX attachments through the appropriate pipeline

How It Works

  1. Text Extraction — native text via PyMuPDF/pdfplumber/python-pptx, with OCR fallback (Tesseract) for posters/talks/scanned patents
  2. Page Rendering — pages rendered to images using PyMuPDF (PDFs) or PowerPoint COM (PPTX) when Vision AI is needed
  3. Vision AI — Claude analyzes page images for figures, chemical structures, and garbled text (smart gating skips this when not needed)
  4. Structuring — extracted content is parsed into logical sections
  5. Summarization — AI generates executive summaries and key findings (always in English)
  6. Quality Scoring — automated scoring flags documents that may need manual review

Documentation

Guide Description
Setup Guide Full installation from scratch — Miniconda, packages, credentials, troubleshooting
Poster Pipeline Deep dive — architecture, processing stages, figure analysis, quality scoring
Patent Pipeline Deep dive — claims parsing, SMILES extraction, text quality repair, semantic classification
Talk Pipeline Deep dive — slide extraction, OCR pre-pass, summary generation, batch processing
Presentation Pipeline Deep dive — PPTX/PDF extraction, classification detection, action items, naming schemes
Outlook Pipeline Deep dive — email threading, incremental processing, attachment routing

Requirements

  • Python 3.11+ (via Conda)
  • Tesseract OCR (poster/talk/patent pipelines — required for poster/talk, optional for patent scanned PDFs)
  • python-pptx, PyMuPDF, pdfplumber (presentation/paper pipelines)
  • python-docx (outlook pipeline — DOCX attachment extraction)
  • pyyaml (outlook pipeline — contact book YAML read/write; usually pre-installed with conda)
  • Microsoft PowerPoint (optional — enables Vision AI for PPTX slide rendering)
  • Microsoft Outlook (outlook pipeline — must be running; no OAuth/admin required)
  • Anthropic API access (Claude Sonnet 4.6)
  • Windows (uses Windows Registry for credential loading; Outlook COM requires Windows)

Note: paper_pipeline.py and ci_pipeline.py do not yet have dedicated README files — see inline --help for their CLI options.

About

AI-powered scientific document processing pipelines (patents, posters, talks)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages