Releases · ggerganov/llama.cpp

09 May 00:59

4426e29

cmake : fix typo (#7151)

Assets 19

cudart-llama-bin-win-cu11.7.1-x64.zip

293 MB 2024-05-09T00:59:42Z
cudart-llama-bin-win-cu12.2.0-x64.zip

413 MB 2024-05-09T00:59:51Z
llama-b2824-bin-macos-arm64.zip

40.3 MB 2024-05-09T01:00:05Z
llama-b2824-bin-macos-x64.zip

36.9 MB 2024-05-09T01:00:07Z
llama-b2824-bin-ubuntu-x64.zip

45.5 MB 2024-05-09T01:00:10Z
llama-b2824-bin-win-arm64-x64.zip

5.98 MB 2024-05-09T01:00:12Z
llama-b2824-bin-win-avx-x64.zip

6.56 MB 2024-05-09T01:00:13Z
llama-b2824-bin-win-avx2-x64.zip

6.53 MB 2024-05-09T01:00:14Z
llama-b2824-bin-win-avx512-x64.zip

6.55 MB 2024-05-09T01:00:15Z
llama-b2824-bin-win-clblast-x64.zip

7.73 MB 2024-05-09T01:00:16Z
Source code (zip)

2024-05-08T23:55:32Z
Source code (tar.gz)

2024-05-08T23:55:32Z

09 May 00:13

github-actions

b2822

bc4bba3

b2822

Introduction of CUDA Graphs to LLama.cpp (#6766)

* DRAFT: Introduction of CUDA Graphs to LLama.cpp

* FIx issues raised in comments

* Tidied to now only use CUDA runtime (not mixed with driver calls)

* disable for multi-gpu and batch size > 1

* Disable CUDA graphs for old GPU arch and with env var

* added missing CUDA_CHECKs

* Addressed comments

* further addressed comments

* limit to GGML_ALLOW_CUDA_GRAPHS defined in llama.cpp cmake

* Added more comprehensive graph node checking

* With mechanism to fall back if graph capture fails

* Revert "With mechanism to fall back if graph capture fails"

This reverts commit eb9f15fb6fcb81384f732c4601a5b25c016a5143.

* Fall back if graph capture fails and address other comments

* - renamed GGML_ALLOW_CUDA_GRAPHS to GGML_CUDA_USE_GRAPHS

- rename env variable to disable CUDA graphs to GGML_CUDA_DISABLE_GRAPHS

- updated Makefile build to enable CUDA graphs

- removed graph capture failure checking in ggml_cuda_error
  using a global variable to track this is not thread safe, but I am also not safistied with checking an error by string
  if this is necessary to workaround some issues with graph capture with eg. cuBLAS, we can pass the ggml_backend_cuda_context to the error checking macro and store the result in the context

- fixed several resource leaks

- fixed issue with zero node graphs

- changed fixed size arrays to vectors

- removed the count of number of evaluations before start capturing, and instead changed the capture mode to relaxed

- removed the check for multiple devices so that it is still possible to use a single device, instead checks for split buffers to disable cuda graphs with -sm row

- changed the op for checking batch size to GGML_OP_ADD, should be more reliable than GGML_OP_SOFT_MAX

- code style fixes

- things to look into
  - VRAM usage of the cudaGraphExec_t, if it is significant we may need to make it optional
  - possibility of using cudaStreamBeginCaptureToGraph to keep track of which ggml graph nodes correspond to which cuda graph nodes

* fix build without cuda graphs

* remove outdated comment

* replace minimum cc value with a constant

---------

Co-authored-by: slaren <[email protected]>

Assets 19

08 May 23:52

github-actions

b2821

c12452c

b2821

JSON: [key] -> .at(key), assert() -> GGML_ASSERT (#7143)

Assets 19

08 May 22:56

github-actions

b2820

9da243b

b2820

Revert "llava : add support for moondream vision language model (#6899)"

This reverts commit 46e12c4692a37bdd31a0432fc5153d7d22bc7f72.

Assets 19

08 May 22:43

github-actions

b2818

26458af

b2818

metal : use `vm_allocate` instead of `posix_memalign` on macOS (#7078)

* fix: use `malloc` instead of `posix_memalign` in `ggml-metal.m` to make it not crash Electron proccesses

* fix: typo

* fix: use `vm_allocate` instead of `posix_memalign`

* fix: don't call `newBufferWithBytesNoCopy` with `NULL` when `ggml_metal_host_malloc` returns `NULL`

* fix: use `vm_allocate` only on macOS

Assets 19

08 May 18:14

github-actions

b2817

83330d8

b2817

main : add --conversation / -cnv flag (#7108)

Assets 19

08 May 17:43

github-actions

b2816

465263d

b2816

sgemm : AVX Q4_0 and Q8_0 (#6891)

* basic avx implementation

* style

* combine denibble with load

* reduce 256 to 128 (and back!) conversions

* sse load

* Update sgemm.cpp

* oops

oops

Assets 19

08 May 17:15

github-actions

b2815

911b390

b2815

server : add_special option for tokenize endpoint (#7059)

Assets 19

08 May 14:52

github-actions

b2813

229ffff

b2813

llama : add BPE pre-tokenization for Qwen2 (#7114)

* Add BPE pre-tokenization for Qwen2.

* minor : fixes

---------

Co-authored-by: Ren Xuancheng <[email protected]>
Co-authored-by: Georgi Gerganov <[email protected]>

Assets 19

08 May 14:50

github-actions

b2812

1fd9c17

b2812

clean up json_value & server_log (#7142)

Assets 19

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Releases: ggerganov/llama.cpp

b2824

b2822

b2821

b2820

b2818

b2817

b2816

b2815

b2813

b2812