Skip to main content

PDF Extractor Pack

TrustedSandbox optionalv1.0.0MITVerified80

by AgentNode · published 5 months ago · toolpack

Extract text, tables, and images from PDF documents.

Parse PDF files to extract structured text, data tables, embedded images, and metadata. Supports OCR for scanned documents via pytesseract.

langchaincrewaigeneric

Quick Start

bash
agentnode install pdf-extractor-pack

Runs in a subprocess with filtered environment by default. Declared permissions are policy-checked, not sandboxed.

Usage

From package
python
from pdf_extractor_pack.tool import run

result = run(
    action="extract_pdf",
    file_path="/tmp/q4-earnings-report.pdf",
    extract=["tables", "text"],
    pages="3-8"
)

print(f"Pages processed: {result['page_count']}")
print(f"Tables found: {len(result['tables'])}")

for i, table in enumerate(result["tables"]):
    print(f"\nTable {i+1} (page {table['page']})")
    print(f"  Columns: {table['headers']}")
    for row in table["rows"][:3]:
        print(f"  {row}")
    print(f"  ... {len(table['rows'])} total rows")

Runs locally on your machine. No execution data is sent to AgentNode. Permissions are checked before execution. Learn how this works

Verification

high confidence80/100✔ Verified
smokeReturned valid result
+25/25
testsTests failed
0/15
importAll tools imported successfully
+15/15
installInstalled in 1.8s
+15/15
contractAll contract checks passed
+10/10
determinismConsistent output across runs (normalized)
+5/5
reliability3/3 runs passed
+10/10

Package installs and imports correctly. runtime checks passed.

install1.8s
import485ms
smoke244ms
tests1.2s

This package was executed and validated by AgentNode before listing. Install, import, and runtime checks passed.

Verified in real_auto mode

Python 3.12.3ffmpegpopplertesseractuv

Last verified 29d ago· Runner v2.0.0

Use this when you need to...

  • Extract financial tables from quarterly earnings PDF reports
  • Pull embedded images and charts from research papers
  • Convert multi-page PDF contracts into structured JSON sections
  • Extract metadata and document properties from legal filings
  • OCR scanned PDF documents that contain no selectable text

README

Version History

Capabilities

pdf_extractionextract_pdftool

Input Schema

{
  "type": "object",
  "required": [
    "file_path"
  ],
  "properties": {
    "pages": {
      "type": "string",
      "default": "all"
    },
    "file_path": {
      "type": "string",
      "default": "/tmp/agentnode_verify/test.pdf",
      "description": "Path to the PDF file"
    },
    "extract_images": {
      "type": "boolean",
      "default": false
    },
    "extract_tables": {
      "type": "boolean",
      "default": true
    }
  }
}

Permissions

Sandbox optionalFrom a trusted publisher — runs on the host by default. You can require isolation with sandbox.host_trust_policy.

Declared by the publisher. Checked before execution by the policy gate.

Networknone
Filesystemtemp
Code Executionnone
Data Accessinput_only
User Approvalnever

Permissions are policy-checked before execution. For trusted and curated packages that run on the host, network and filesystem access are policy-checked but not OS-sandboxed. When runtime isolation is required for untrusted/community code, AgentNode uses sandbox-or-fail-closed if the required container runtime and pinned image are available. Learn more

Privacy

All tool execution happens locally on your machine. AgentNode never receives:

  • • Tool inputs or outputs
  • • Execution logs
  • • Data your agent processes

Only install events and search queries are sent to the registry.

bash
agentnode install pdf-extractor-pack

Files (3)

License

MIT

Stats

Downloads0
Installs0
Versionv1.0.0
Published3/16/2026
Channelstable
Typetoolpack
Entrypointpdf_extractor_pack.tool

Compatibility

Frameworks

langchaincrewaigeneric

Runtime

python

Python Version

>=3.10

Trust & Security

PublisherTrusted
SignatureNone
ProvenanceNone
Security Issues0

Publisher

A

AgentNode

@agentnode