README.md

Book Agent 0.1.0 · 本版随附原文,按章节提供导览;完整原文可在文末展开。文内本机路径属于示例,请替换为你的实际路径。

本版本其他文档与许可
# Original BGE-M3 query model bundle

This optional directory supplies the inherited `BAAI/bge-m3` query encoder for Book Agent. It is separate from book packages, document vectors and the portable executable. Share this directory among books with the same audited contract. No API key is required for CPU queries.

Register after moving

Install Book Agent's optional `local-query` Python dependencies, put this directory at a chosen location, and register its actual path:

```shell
python -m pip install "prebuilt-book-agent[local-query]"
book-agent install-query-model BOOK_ID --model-dir ./BAAI-bge-m3
book-agent configure BOOK_ID --vector-search offline --query-model-path ./BAAI-bge-m3
```

Registration without `--download` verifies the local files and portable `book-agent-query-contract.json` receipt. It never downloads. The receipt contains relative artifact names and hashes, with no personal paths or credentials. Keep it with the files. The default vector mode is `none`; fulltext and source retrieval work without activating this model.

返回章节目录

Pinned identity

- Model: [BAAI/bge-m3](https://huggingface.co/BAAI/bge-m3)
- Original tokenizer/configuration revision: `5617a9f61b028005a4858fdac845db406aefb181`
- Safetensors conversion revision: `9a0624b896d81da7492a910ffa53731274b6cf3d`
- Recipe: `bge-m3-dense-cls-l2-float32-v1`
- Output: 1024 dimensions, original fast XLM-Roberta tokenizer, raw question without instruction/prefix, CLS pooling, L2 normalization, CPU float32

| Artifact | Bytes | SHA-256 |
| --- | ---: | --- |
| `config.json` | 687 | `26159e7ad065073448460117eb24b7a4572f6f4e78eadff65dc0a11c052449fa` |
| `tokenizer_config.json` | 444 | `a62b2b6784f990259fddef5f16388693a8043be4f69179e6a5257eeb3f9abac4` |
| `special_tokens_map.json` | 964 | `8c785abebea9ae3257b61681b4e6fd8365ceafde980c21970d001e834cf10835` |
| `tokenizer.json` | 17,098,108 | `21106b6d7dab2952c1d496fb21d5dc9db75c28ed361a05f5020bbba27810dd08` |
| `sentencepiece.bpe.model` | 5,069,051 | `cfc8146abe2a0488e9e2a0c56de7952f7c11ab059eca145a0a727afce0db2865` |
| `model.safetensors` | 2,271,064,456 | `993b2248881724788dcab8c644a91dfd63584b6e5604ff2037cb5541e1e38e7e` |

The original upstream revision has `pytorch_model.bin`. This bundle uses the public Hugging Face [conversion bot PR #130](https://huggingface.co/BAAI/bge-m3/discussions/130) safetensors commit, a direct child of that original revision. The conversion PR remains unmerged. Book Agent does not download/load the unsafe binary and disables remote code and implicit model downloads during search.

Identity, artifact hashes, tokenizer and CLS/L2 recipe are fixed. Numerical parity with the original unpinned SiliconFlow hosted encoder has not been independently verified. Runtime output reports `provider_numeric_parity_verified: false`. The bundle does not claim that hosted-service floating-point results are identical.

返回章节目录

CPU budget

The query adapter uses one encoder layer at a time in CPU float32 RAM and reads only the question's embedding rows. It does not retain the whole 2.27 GB checkpoint in memory or duplicate weights. Each question rereads encoder parameters, so slow storage affects latency. Defaults are two CPU inference threads and 512 tokenizer tokens including special tokens. Configuration allows 1–4 threads and 1–8192 tokens. The Core question limit is 8,000 characters; the lower-level encoder independently caps 16,000. Oversized questions are rejected before forward, without truncation. The weight file remains required on disk, plus dependencies and activation memory. Documents are never re-embedded.

See [MODEL-LICENSE.md](MODEL-LICENSE.md) and the original [UPSTREAM-LICENSE.txt](UPSTREAM-LICENSE.txt) for upstream provenance and license information.

返回章节目录

查看完整原文(逐字保留)
# Original BGE-M3 query model bundle

This optional directory supplies the inherited `BAAI/bge-m3` query encoder for Book Agent. It is separate from book packages, document vectors and the portable executable. Share this directory among books with the same audited contract. No API key is required for CPU queries.

## Register after moving

Install Book Agent's optional `local-query` Python dependencies, put this directory at a chosen location, and register its actual path:

```shell
python -m pip install "prebuilt-book-agent[local-query]"
book-agent install-query-model BOOK_ID --model-dir ./BAAI-bge-m3
book-agent configure BOOK_ID --vector-search offline --query-model-path ./BAAI-bge-m3
```

Registration without `--download` verifies the local files and portable `book-agent-query-contract.json` receipt. It never downloads. The receipt contains relative artifact names and hashes, with no personal paths or credentials. Keep it with the files. The default vector mode is `none`; fulltext and source retrieval work without activating this model.

## Pinned identity

- Model: [BAAI/bge-m3](https://huggingface.co/BAAI/bge-m3)
- Original tokenizer/configuration revision: `5617a9f61b028005a4858fdac845db406aefb181`
- Safetensors conversion revision: `9a0624b896d81da7492a910ffa53731274b6cf3d`
- Recipe: `bge-m3-dense-cls-l2-float32-v1`
- Output: 1024 dimensions, original fast XLM-Roberta tokenizer, raw question without instruction/prefix, CLS pooling, L2 normalization, CPU float32

| Artifact | Bytes | SHA-256 |
| --- | ---: | --- |
| `config.json` | 687 | `26159e7ad065073448460117eb24b7a4572f6f4e78eadff65dc0a11c052449fa` |
| `tokenizer_config.json` | 444 | `a62b2b6784f990259fddef5f16388693a8043be4f69179e6a5257eeb3f9abac4` |
| `special_tokens_map.json` | 964 | `8c785abebea9ae3257b61681b4e6fd8365ceafde980c21970d001e834cf10835` |
| `tokenizer.json` | 17,098,108 | `21106b6d7dab2952c1d496fb21d5dc9db75c28ed361a05f5020bbba27810dd08` |
| `sentencepiece.bpe.model` | 5,069,051 | `cfc8146abe2a0488e9e2a0c56de7952f7c11ab059eca145a0a727afce0db2865` |
| `model.safetensors` | 2,271,064,456 | `993b2248881724788dcab8c644a91dfd63584b6e5604ff2037cb5541e1e38e7e` |

The original upstream revision has `pytorch_model.bin`. This bundle uses the public Hugging Face [conversion bot PR #130](https://huggingface.co/BAAI/bge-m3/discussions/130) safetensors commit, a direct child of that original revision. The conversion PR remains unmerged. Book Agent does not download/load the unsafe binary and disables remote code and implicit model downloads during search.

Identity, artifact hashes, tokenizer and CLS/L2 recipe are fixed. Numerical parity with the original unpinned SiliconFlow hosted encoder has not been independently verified. Runtime output reports `provider_numeric_parity_verified: false`. The bundle does not claim that hosted-service floating-point results are identical.

## CPU budget

The query adapter uses one encoder layer at a time in CPU float32 RAM and reads only the question's embedding rows. It does not retain the whole 2.27 GB checkpoint in memory or duplicate weights. Each question rereads encoder parameters, so slow storage affects latency. Defaults are two CPU inference threads and 512 tokenizer tokens including special tokens. Configuration allows 1–4 threads and 1–8192 tokens. The Core question limit is 8,000 characters; the lower-level encoder independently caps 16,000. Oversized questions are rejected before forward, without truncation. The weight file remains required on disk, plus dependencies and activation memory. Documents are never re-embedded.

See [MODEL-LICENSE.md](MODEL-LICENSE.md) and the original [UPSTREAM-LICENSE.txt](UPSTREAM-LICENSE.txt) for upstream provenance and license information.

原文 SHA-256:6141278fbe1849c5d287cd5077166b59db9d30f2e7d5f0086c218a2360ac3fa6