使用 xberg Go 绑定从远程文本文档 URL 提取内容:URI 输入与 URL 提取配置实战
后端AI 应用NLP【免费下载链接】xbergPolyglot document intelligence with a Rust core: extract text, metadata, images, tables, and structured data from 106 formats across 140 file extensions, plus code intelligence for 371 languages. Fifteen bindings, with CLI, REST API, and MCP server.项目地址https://gitcode.com/gh_mirrors/kr/xberg点击查看免费下载本篇指南聚焦 xberg 项目的 Go 语言绑定github.com/xberg-io/xberg/packages/go中一项非常实用的能力以 HTTP(S) 远程文本文档 URL 作为输入通过ExtractInput{Kind: uri}与ExtractionConfig{URL.Mode: document}组合一次性取回文档正文与远端 URL 统计信息。读完本文你将掌握 xberg Go API 中 URI 输入、URL 提取模式的完整配置方式并能结合仓库内的 e2e 测试与 fixture 数据在自有 Go 服务中直接落地拉取远程文档并提取文本的流程。一、场景与核心概念为什么用kind: uri而不是kind: bytesxberg 的Extract统一入口支持两类输入源定义在 packages/go/binding.go 的ExtractInputKind枚举中枚举值JSON 取值适用场景ExtractInputKindBytesbytes内存中已有的原始字节配合mime_type/filename提示进行类型识别ExtractInputKindURIuri本地文件路径、file://URI或 HTTP(S) 远程 URL当文档不在本地磁盘、而是托管在远端服务器文本文件、日志、配置文件、发布说明等时uri是最自然的输入方式只需把 URL 字符串交给提取引擎xberg 会自动完成下载、MIME 识别、内容解析与文本提取无需在客户端先行读取字节。ExtractInput结构体完整字段如下源码位置type ExtractInput struct { // Source kind. bytes requires bytes; uri requires uri. Kind *ExtractInputKind json:kind,omitempty // Raw bytes for kind bytes. Bytes []byte json:bytes,omitempty // Local path, file:// URI, or HTTP(S) URL for kind uri. URI *string json:uri,omitempty // MIME type hint. MimeType *string json:mime_type,omitempty // Filename hint used for MIME detection and metadata. Filename *string json:filename,omitempty // Per-input extraction overrides. Config *FileExtractionConfig json:config,omitempty }远程 URL 场景下只需填充Kind与URI两个字段其余保持零值即可。二、完整示例从远程文本文档 URL 提取内容原文档docs-site/src/snippets-generated/go/url/url_remote_text_document.md给出了一个可直接运行的 Go 示例其核心调用链如下我们补上逐行注释package main import ( fmt xberg github.com/xberg-io/xberg/packages/go ) // ptr 是 xberg Go 绑定提供的小工具 // 用于为可选字段构造指针与绑定中的 xberg.Ptr 等价。 func ptrT any *T { return value } func main() { // 1. 构造输入声明输入类型为 URI并指定远程文档地址 input : xberg.ExtractInput{ Kind: ptr(xberg.ExtractInputKindURI), URI: ptr(https://example.com), } // 2. 构造提取配置URL 部分设置为 document 模式 // 表示把该 URI 当作单个远程文档/页面处理 config : xberg.ExtractionConfig{ URL: xberg.URLExtractionConfig{ Mode: ptr(xberg.URLExtractionModeDocument), }, } // 3. 执行提取 result, err : xberg.Extract(input, config) if err ! nil { panic(err) } // 4. 打印第一个结果的正文内容 fmt.Printf(%v\n, result.Results[0].Content) // 5. 打印摘要中的远端 URL 计数 fmt.Printf(%v\n, result.Summary.RemoteUrls) }这段代码对应仓库中同名 e2e 测试 Test_UrlRemoteTextDocument测试通过 mock 服务器返回一段text/plain; charsetutf-8文本内容为Remote document hello from Xberg URL e2e.断言提取结果包含该文本、且Summary.RemoteUrls 1。三、配置拆解URLExtractionConfig与mode: document示例中最关键的一行是config.URL xberg.URLExtractionConfig{Mode: ptr(xberg.URLExtractionModeDocument)}。URLExtractionMode定义了三种 URL 处理策略源码位置模式常量JSON 取值行为URLExtractionModeAutoauto拉取 HTTP(S) 资源后自动分类根据响应内容判断是单文档还是需要爬取URLExtractionModeDocumentdocument把 URI 当作单个远程文档/页面直接下载并提取——本篇主题URLExtractionModeCrawlcrawl以种子 URI 为起点爬取并提取发现的所有页面/文档document模式适用于明确知道目标是一个具体文档的场景服务器返回什么内容就按什么格式解析。对于返回text/plain的文本文件提取引擎会直接产出纯文本正文若 URL 指向 PDF、DOCX 等二进制文档也会走对应的格式解析管线xberg 核心支持 106 种格式、140 种文件扩展名。URLExtractionConfig还包含其他可选字段源码位置可结合具体需求继续配置type URLExtractionConfig struct { // URL extraction mode. Mode *URLExtractionMode json:mode,omitempty // Crawlberg crawl configuration used for HTTP(S) URL extraction. Crawl *CrawlConfig json:crawl,omitempty // Optional regex filter for document-discovered URLs. DocumentURLPattern *string json:document_url_pattern,omitempty // Maximum URLs to follow per extraction result. MaxDocumentUrlsPerResult *uint32 json:max_document_urls_per_result,omitempty // Maximum URLs followed across the whole extraction call. MaxTotalUrls *uint32 json:max_total_urls,omitempty // Allow bare local filesystem path inputs. AllowLocalFileInputs *bool json:allow_local_file_inputs,omitempty // Allow local file:// URI inputs. AllowFileUris *bool json:allow_file_uris,omitempty }其中CrawlCrawlConfig与document模式并非互斥从单文档结果中发现的链接仍可按需跟随。仓库 fixture url_recursive_document_urls.json 演示了mode: document配合crawl子配置的用法{ url: { mode: document, crawl: { follow_document_urls: true, document_url_depth: 1, respect_robots_txt: false } } }该 fixture 用 mock 服务器返回一个带a href/linked.txt链接的 HTML 页面随后递归拉取/linked.txt并断言第二个结果包含目标文本——这证明即使在document模式下xberg 仍支持从文档内发现的 URL 继续提取的递归能力。四、结果结构正文、摘要与远端 URL 统计xberg.Extract返回*ExtractionResult其字段定义见 packages/go/binding.gotype ExtractionResult struct { // Extracted documents in discovery order. Results []ExtractedDocument json:results,omitempty // Non-fatal per-input errors. Errors []ExtractionErrorItem json:errors,omitempty // Aggregate counts for the operation. Summary *ExtractionSummary json:summary,omitempty // Final URLs reached after redirects during URL ingestion. CrawlFinalUrls []string json:crawl_final_urls,omitempty // Total redirects followed while fetching or crawling URLs. CrawlRedirectCount uint json:crawl_redirect_count // Unique normalized URLs discovered by crawls. CrawlUniqueNormalizedUrls []string json:crawl_unique_normalized_urls,omitempty } type ExtractionSummary struct { // Number of inputs submitted by the caller. Inputs uint json:inputs // Number of extraction results produced. Results uint json:results // Number of per-input errors. Errors uint json:errors // Number of URI inputs that resolved to remote HTTP(S) URLs. RemoteUrls uint json:remote_urls // Number of HTML pages crawled or scraped. PagesCrawled uint json:pages_crawled // Number of downloaded non-HTML documents extracted from URLs. DocumentsDownloaded uint json:documents_downloaded }示例中打印的两项含义明确result.Results[0].Content第一个也是本例中唯一一个提取结果的正文字段result.Summary.RemoteUrls本次调用中解析为远端 HTTP(S) URL 的输入数量。单个 URL 输入成功提取后该值为1这正是 e2e 测试断言equals summary.remote_urls 1的语义对应 fixture url_remote_text_document.json 的断言配置。五、配套能力批量混合输入与递归 URL 提取单 URL 提取是基础xberg 还提供同族扩展能力均以URLExtractionConfig为配置中心批量混合输入extract_batch。fixture url_batch_mixed_inputs.json 演示了URL 输入 字节输入共享一个输出信封的用法inputs数组里既有{kind:uri,uri:$mock_url}也有带内联字节、mime_type与filename的 bytes 输入二者共用{url:{mode:document}}配置最终results数组按输入顺序返回两份内容。适合远端文档 本地内联文件混合批处理的场景。递归文档 URL 提取crawl.follow_document_urls。如第三节所述通过crawl子配置可控制是否跟随文档内发现的链接、递归深度以及是否遵守robots.txt见 url_recursive_document_urls.json。网页/文档模式切换。document模式也适用于普通网页——同主题的 url_html_page_extract.md 展示了用相同配置提取https://example.com页面内容的写法区别仅在于最终打印result.Results全部结果而非单个字段。六、错误处理与工程化要点错误语义xberg.Extract返回的 error 由 FFI 层映射而来Go 绑定中定义了ErrXbergIo、ErrParse、ErrTimeout、ErrUnsupportedFormat、ErrSecurity等哨兵错误packages/go/binding.go可用errors.Is精确匹配nativeError同时保留原生层细节消息如耗时、限制与计数。指针构造xberg Go API 的字段大量使用指针表达可选建议统一使用绑定自带的xberg.Ptr[T]或示例中的局部ptr泛型函数构造避免手写地址取用。本地与远端统一入口uri同时覆盖本地路径、file://与 HTTP(S) URLURLExtractionConfig.AllowLocalFileInputs/AllowFileUris可控制是否放行本地输入便于在多租户/服务端场景收紧安全边界。远端 URL 统计判断是否真的走了网络提取可依赖Summary.RemoteUrls配合CrawlRedirectCount与CrawlFinalUrls可追踪重定向链路适合做链路观测与排障。结语从远程文本文档 URL 提取内容是 xberg 统一Extract入口中最轻量、最常用的路径之一ExtractInput{Kind: uri, URI: url}声明输入URLExtractionConfig{Mode: document}声明处理策略一次调用即可拿到正文与Summary.RemoteUrls统计。结合仓库中 e2e 测试 与 fixtures/url 目录下的真实断言数据你可以快速验证并扩展出批量、递归、网页抓取等更复杂的远程内容提取管线。赞分享后端AI 应用NLP【免费下载链接】xbergPolyglot document intelligence with a Rust core: extract text, metadata, images, tables, and structured data from 106 formats across 140 file extensions, plus code intelligence for 371 languages. Fifteen bindings, with CLI, REST API, and MCP server.项目地址https://gitcode.com/gh_mirrors/kr/xberg点击查看免费下载相关推荐Xberg Dart 绑定 URI 提取实战从 URL 与本地路径抽取文档内容Xberg Dart 绑定 URI 提取实战从 URL 与本地路径抽取文档内容 本篇技术指南聚焦 Xberg 项目中 Dart 语言绑定的 URI 提取 AP后端AI 应用NLPXberg C 绑定 URI 提取实战使用 ExtractAsync 从 URL 抽取 PDF 内容Xberg C 绑定 URI 提取实战使用 ExtractAsync 从 URL 抽取 PDF 内容 本篇技术指南聚焦 Xberg 官方 C 绑定中的 URI后端AI 应用NLP使用 Xberg Dart 绑定通过 URL 提取远程文本文档url.modedocument 实战指南使用 Xberg Dart 绑定通过 URL 提取远程文本文档url.modedocument 实战指南 本篇技术指南围绕 Xberg 仓库中 Dart 语后端AI 应用NLP上一篇lightweight-charts 插件渲染入门深入理解 CanvasRenderingTarget2D 与 Bitmap / Media 双坐标系下一篇告别调试烦恼DreamBerd问号语法让错误处理更轻松创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考