Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
一个用Rust编写的PDF快速解析库,能判断PDF是文本型还是扫描型,提取文字并转为Markdown,支持Python、Node.js和浏览器WebAssembly绑定,无OCR依赖。