convert-llama2c-to-ggml : enable conversion of GQA models (ggerganov#…

…6237) * convert-llama2c-to-ggml: enable conversion of multiqueries, ggerganov#5608 * add test in build action * Update build.yml * Update build.yml * Update build.yml * gg patch
hodlen · Apr 3, 2024 · 2cd14a4 · 2cd14a4
1 parent 6d9e216
commit 2cd14a4
Show file tree

Hide file tree

Showing 3 changed files with 194 additions and 208 deletions.
diff --git a/.github/workflows/build.yml b/.github/workflows/build.yml
@@ -225,6 +225,17 @@ jobs:
  cd build
  ctest -L main --verbose --timeout 900
 
+ - name: Test llama2c conversion
+ id: llama2c_test
+ run: |
+ cd build
+ echo "Fetch tokenizer"
+ wget https://huggingface.co/karpathy/tinyllamas/resolve/main/stories260K/tok512.bin
+ echo "Fetch llama2c model"
+ wget https://huggingface.co/karpathy/tinyllamas/resolve/main/stories260K/stories260K.bin
+ ./bin/convert-llama2c-to-ggml --copy-vocab-from-model ./tok512.bin --llama2c-model stories260K.bin --llama2c-output-model stories260K.gguf
+ ./bin/main -m stories260K.gguf -p "One day, Lily met a Shoggoth" -n 500 -c 256
+
 # ubuntu-latest-cmake-sanitizer:
 # runs-on: ubuntu-latest
 #

diff --git a/examples/convert-llama2c-to-ggml/README.md b/examples/convert-llama2c-to-ggml/README.md
@@ -21,6 +21,8 @@ An example command using a model from [karpathy/tinyllamas](https://huggingface.
 
 `$ ./convert-llama2c-to-ggml --copy-vocab-from-model llama-2-7b-chat.gguf.q2_K.bin --llama2c-model stories42M.bin --llama2c-output-model stories42M.gguf.bin`
 
+Note: The vocabulary for `stories260K.bin` should be its own tokenizer `tok512.bin` found in [karpathy/tinyllamas/stories260K](https://huggingface.co/karpathy/tinyllamas/tree/main/stories260K).
+
 Now you can use the model with a command like:
 
 `$ ./main -m stories42M.gguf.bin -p "One day, Lily met a Shoggoth" -n 500 -c 256`