@huggingface/gguf

一个适用于远程托管文件的 GGUF 解析器。

规范

规范: https://github.com/ggerganov/ggml/blob/master/docs/gguf.md

参考实现 (Python): https://github.com/ggerganov/llama.cpp/blob/master/gguf-py/gguf/gguf_reader.py

安装

npm install @huggingface/gguf

使用方法

基本用法

import { GGMLQuantizationType, gguf } from "@huggingface/gguf";

// remote GGUF file from https://huggingface.co/TheBloke/Llama-2-7B-Chat-GGUF
const URL_LLAMA = "https://huggingface.co/TheBloke/Llama-2-7B-Chat-GGUF/resolve/191239b/llama-2-7b-chat.Q2_K.gguf";

const { metadata, tensorInfos } = await gguf(URL_LLAMA);

console.log(metadata);
// {
//     version: 2,
//     tensor_count: 291n,
//     kv_count: 19n,
//     "general.architecture": "llama",
//     "general.file_type": 10,
//     "general.name": "LLaMA v2",
//     ...
// }

console.log(tensorInfos);
// [
//     {
//         name: "token_embd.weight",
//         shape: [4096n, 32000n],
//         dtype: GGMLQuantizationType.Q2_K,
//     },

//     ... ,

//     {
//         name: "output_norm.weight",
//         shape: [4096n],
//         dtype: GGMLQuantizationType.F32,
//     }
// ]

读取本地文件

// Reading a local file. (Not supported on browser)
const { metadata, tensorInfos } = await gguf(
  './my_model.gguf',
  { allowLocalFile: true },
);

严格类型

默认情况下，metadata 中的已知字段是类型化的。这包括在 llama.cpp、whisper.cpp 和 ggml 中找到的各种字段。

const { metadata, tensorInfos } = await gguf(URL_MODEL);

// Type check for model architecture at runtime
if (metadata["general.architecture"] === "llama") {

  // "llama.attention.head_count" is a valid key for llama architecture, this is typed as a number
  console.log(model["llama.attention.head_count"]);

  // "mamba.ssm.conv_kernel" is an invalid key, because it requires model architecture to be mamba
  console.log(model["mamba.ssm.conv_kernel"]); // error
}

禁用严格类型

由于 GGUF 格式可以用于存储张量，我们可以在技术上将其用于其他用途。例如，存储控制向量、lora 权重等。

如果您想使用自己的 GGUF 元数据结构，您可以通过将解析输出转换为 GGUFParseOutput<{ strict: false }> 来禁用严格类型。

const { metadata, tensorInfos }: GGUFParseOutput<{ strict: false }> = await gguf(URL_LLAMA);

命令行界面

此包提供了一个与 gguf_dump.py 脚本等效的 CLI。您可以使用此命令转储 GGUF 元数据和张量列表

npx @huggingface/gguf my_model.gguf

# or, with a remote GGUF file:
# npx @huggingface/gguf https://huggingface.co/bartowski/Llama-3.2-1B-Instruct-GGUF/resolve/main/Llama-3.2-1B-Instruct-Q4_K_M.gguf

输出示例

* Dumping 36 key/value pair(s)
  Idx | Count  | Value                                                                            
  ----|--------|----------------------------------------------------------------------------------
    1 |      1 | version = 3                                                                      
    2 |      1 | tensor_count = 292                                                               
    3 |      1 | kv_count = 33                                                                    
    4 |      1 | general.architecture = "llama"                                                   
    5 |      1 | general.type = "model"                                                           
    6 |      1 | general.name = "Meta Llama 3.1 8B Instruct"                                      
    7 |      1 | general.finetune = "Instruct"                                                    
    8 |      1 | general.basename = "Meta-Llama-3.1"                                                   

[truncated]

* Dumping 292 tensor(s)
  Idx | Num Elements | Shape                          | Data Type | Name                     
  ----|--------------|--------------------------------|-----------|--------------------------
    1 |           64 |     64,      1,      1,      1 | F32       | rope_freqs.weight        
    2 |    525336576 |   4096, 128256,      1,      1 | Q4_K      | token_embd.weight        
    3 |         4096 |   4096,      1,      1,      1 | F32       | blk.0.attn_norm.weight   
    4 |     58720256 |  14336,   4096,      1,      1 | Q6_K      | blk.0.ffn_down.weight

[truncated]

或者，您可以将此包全局安装，这将提供 gguf-view 命令

npm i -g @huggingface/gguf
gguf-view my_model.gguf

Hugging Face Hub

Hub 支持所有文件格式，并为 GGUF 格式内置了功能。

更多信息请访问： https://huggingface.co/docs/hub/gguf。

致谢与灵感

https://github.com/hyparam/hyllama 作者 @platypii (MIT 许可证)
https://github.com/ahoylabs/gguf.js 作者 @biw @dkogut1996 @spencekim (MIT 许可证)

🔥❤️

< > 在 GitHub 上更新

Huggingface.js