Twinkle HubTwinkle Hub
登入

📌 2026-08-17 新增:⚖️ 台灣法條時光機 — 歷史法條原文 · 修法沿革 · 全文檢索 · 條文關聯圖(935 個 dataset · 35,454+ 筆 row)

查看完整 Changelog →

抽 PDF 內文

tw_extract_pdf_text

通用MIT

提取 PDF 內文(pymupdf-based, born-digital only)。

Args:
    source: PDF URL 或 base64-encoded PDF 內容。
    max_pages: 最多處理幾頁(None = 全部)。

Returns: text + per-page list + is_scanned flag。

**行為說明(重要,client agents 請注意)**:
- 對 born-digital PDF(含可選取文字):正常回 text + pages
- 對 scanned / image-only PDF:**這不是錯誤**,會優雅降級回
  `{is_scanned: true, reason: "OCR required, ..."}`,無 `error` 欄位、
  也不丟例外。掃描件 OCR 屬高階方案 call_agent 範圍,本 tool 不處理。
- 下游 agent 看到 `is_scanned: true` 應該當作「PDF 不適合此 tool」而非
  「呼叫失敗」,可改走 OCR 或請使用者重供 born-digital 版本。

輸入 schema

{
  "properties": {
    "source": {
      "title": "Source",
      "type": "string"
    },
    "max_pages": {
      "anyOf": [
        {
          "type": "integer"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "title": "Max Pages"
    }
  },
  "required": [
    "source"
  ],
  "title": "extract_pdf_textArguments",
  "type": "object"
}

呼叫​方式

配置​好 Twinkle Hub 的 MCP client 後,​agent 會​看到 tw_extract_pdf_text。直接讓它呼叫​即可,​例如:

# Ask Claude / any MCP client:
請用 tw_extract_pdf_text 處理 "…"。

# It will call:
tw_extract_pdf_text(input="…")

尚未​設定 Client?

Claude Desktop 約 3 分鐘即可完成​設定 — 下載 .mcpb 檔案並雙擊​即可,或參閱​文件中的各 Client 安裝​步驟。

查看使用者設定​文件