[pull] master from ggerganov:master #119

pull · 2024-06-20T09:38:16Z

See Commits and Changes for more details.

Can you help keep this open source service alive? 💖 Please sponsor : )

@slaren

* un-ignore `build-info.cmake` and `build-info.sh` I am assuming that ignoring them was unintentional. If they are ignored, some tools, like cargo, will consider the files inexistent, even if they're comitted, for the purpose of publishing. This leads to the build failing in such cases. * un-ignore `build-info.cpp.in` For the same reason as the previous two files. * Reorganize `.gitignore` * Add exceptions for files mentioned by @slaren I did leave .clang-tidy since it was explicitly ignored before. * Add comments for organization * Sort some lines for pretty * Test with `make` and `cmake` builds to ensure no build artifacts might be comitted * Remove `.clang-tidy` from `.gitignore` Per comment by @ggerganov * Remove `IDEWorkspaceChecks.plist` from root-level `.gitignore`

Currently the Metal backend does not support BF16. `ggml_metal_supports_op` was returning true in these cases, leading to a crash with models converted with `--leave-output-tensor`. This commit checks if the first few sources types are BF16 and returns false if that's the case.

* CUDA: stream-k decomposition for MMQ * fix undefined memory reads for small matrices

* add sycl preset * fix debug link error. fix windows crash * update README

* common: fix warning * Update common/common.cpp Co-authored-by: slaren <[email protected]> --------- Co-authored-by: slaren <[email protected]>

)

* create append_pooling operation; allow to specify attention_type; add last token pooling; update examples * find result_norm/result_embd tensors properly; update output allocation logic * only use embd output for pooling_type NONE * get rid of old causal_attn accessor * take out attention_type; add in llama_set_embeddings * bypass logits when doing non-NONE pooling

ggml-ci

* initial iq4_xs * fix ci * iq4_nl * iq1_m * iq1_s * iq2_xxs * iq3_xxs * iq2_s * iq2_xs * iq3_s before sllv * iq3_s * iq3_s small fix * iq3_s sllv can be safely replaced with sse multiply

mdegans and others added 3 commits June 19, 2024 22:10

server : fix smart slot selection (#8020)

ba58993

github-actions bot added examples server labels Jun 20, 2024

pull bot added ⤵️ pull and removed examples server labels Jun 20, 2024

CUDA: stream-k decomposition for MMQ (#8018)

d50f889

* CUDA: stream-k decomposition for MMQ * fix undefined memory reads for small matrices

github-actions bot added examples server ggml Nvidia GPU labels Jun 20, 2024

[SYCL] Fix windows build and inference (#8003)

de391e4

* add sycl preset * fix debug link error. fix windows crash * update README

github-actions bot added SYCL build labels Jun 20, 2024

JohannesGaessler and others added 2 commits June 20, 2024 16:40

common: fix warning (#8036)

abd894a

* common: fix warning * Update common/common.cpp Co-authored-by: slaren <[email protected]> --------- Co-authored-by: slaren <[email protected]>

convert-hf : Fix the encoding in the convert-hf-to-gguf-update.py (#8040

17b291a

)

github-actions bot added the python label Jun 20, 2024

hamdoudhakem and others added 5 commits June 20, 2024 22:01

requirements : Bump torch and numpy for python3.12 (#8041)

b1ef562

swiftui : enable stream updating (#7754)

0e64591

llama : optimize long word tokenization with WPM (#8034)

a927b0f

ggml-ci

ggml : AVX IQ quants (#7845)

7d5e877

* initial iq4_xs * fix ci * iq4_nl * iq1_m * iq1_s * iq2_xxs * iq3_xxs * iq2_s * iq2_xs * iq3_s before sllv * iq3_s * iq3_s small fix * iq3_s sllv can be safely replaced with sse multiply

teleprint-me closed this Jun 21, 2024

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[pull] master from ggerganov:master #119

[pull] master from ggerganov:master #119

pull bot commented Jun 20, 2024 •

edited

Loading

[pull] master from ggerganov:master #119

[pull] master from ggerganov:master #119

Conversation

pull bot commented Jun 20, 2024 • edited Loading

pull bot commented Jun 20, 2024 •

edited

Loading