PDF Skill for AI Agents: Read, Merge, Split, OCR & Fill Any PDF

Anthropic's official PDF skill. Lets AI agents extract text and tables, merge and split files, rotate pages, add watermarks, fill forms, encrypt and OCR scanned PDFs.

PDF is the file format AI workflows get stuck on most — designed for print, not for machines. Anthropic's official PDF Skill fixes exactly that: it lets Claude Code and compatible agents read, write, edit, merge, split, fill, encrypt, and OCR PDFs the way a human would.

What it covers: the full PDF lifecycle

This skill handles nearly every high-frequency PDF operation:

  • Reading & extraction: pull text and tables from PDFs (pdfplumber gives precise table structure), read metadata (title/author/subject)

  • Merge & split: combine multiple PDFs into one, or split a file into single pages or page ranges

  • Page operations: rotate pages, add watermarks, reorder pages

  • Creation & forms: generate PDFs from scratch, fill PDF forms (the FORMS module handles form fields specifically)

  • Security: encrypt and decrypt PDFs to protect sensitive documents

  • OCR: recognize text in scans so unsearchable PDFs become searchable

Under the hood: Python ecosystem

The skill combines battle-tested Python libraries: pypdf for core read/write and page operations, pdfplumber for high-precision text/table extraction, and a Tesseract pipeline for OCR. The agent picks the right combination for each task automatically — no user intervention needed.

Typical use cases

  • Contract review: read a 50-page contract PDF, extract key clauses, dates, and amounts, and output a structured summary

  • Report assembly: merge weekly report PDFs into a monthly report, watermark it, and distribute

  • Form automation: batch-fill PDF forms (invoices, applications) to replace manual data entry

  • Scan digitization: OCR paper documents into searchable text and feed them into downstream pipelines

How to use & download

Claude Code users: drop the skills/pdf directory into your project's .claude/skills/ folder. Once enabled, just describe the task in natural language ("turn this PDF into text") and the agent handles the rest. Full source and usage notes: official GitHub repo.

Other agent users: the skill follows the standard SKILL.md convention (name/description metadata + operation guide), compatible with any toolchain that supports the Agent Skills standard. Official docs: Anthropic Agent Skills documentation.

Related resources

Combine it with the rest of the document-processing quartet: Word skill (docx), Excel skill (xlsx), and PowerPoint skill (pptx). For general-purpose document OCR, pair it with PaddleOCR.