Fattn logit softcap #244

Nexesenex · 2024-07-17T19:38:56Z

No description provided.

* lora: load to devide buft * add patch tensor function * correct tensor patch * llama_lora_adapter_apply * correct ggml_backend_tensor_copy * add llm_build_mm * fix auto merge * update based on review comments * add convert script * no more transpose A * add f16 convert * add metadata check * add sanity check * fix ftype * add requirements * fix requirements * fix outfile * conversion: only allow selected models * fix types * cuda : do not use dmmv if the tensor does not have enough cols * llama : lora fixes * do not disable mmap with lora Co-authored-by: slaren <[email protected]> * llm_build_lora_mm_id * convert_lora : MoE LoRA conversion support * convert_lora : prefer safetensors, similarly to convert_hf * convert_hf : simplify modify_tensors for InternLM2 * convert_lora : lazy conversion * llama : load and use alpha from LoRA adapters * llama : use llm_build_lora_mm in most model graphs * auto scale * Revert "auto scale" This reverts commit 42415a4. * remove redundant params * Apply suggestions from code review Co-authored-by: slaren <[email protected]> * change kv metadata * move add_type to __init__ * convert_hf : move add_type to main() * convert_lora : use the GGUFWriter from Model instead of overwriting it --------- Co-authored-by: slaren <[email protected]> Co-authored-by: Francis Couture-Harpin <[email protected]>

* convert_hf : faster lazy safetensors This makes '--dry-run' much, much faster. * convert_hf : fix memory leak in lazy MoE conversion The '_lazy' queue was sometimes self-referential, which caused reference cycles of objects old enough to avoid garbage collection until potential memory exhaustion.

The --help option on export-lora isn't accepted as valid. The help still gets displayed by default, but the script exits with an error message and nonzero status.

…nov#8491) * Update clib.json to point to Cyan4973 original xxhash Convinced Cyan4973 to add clib.json directly to his repo, so can now point the clib package directly to him now. Previously pointed to my fork with the clib.json package metadata Cyan4973/xxHash#954 * gguf-hash: readme update to point to Cyan4973 xxHash repo [no ci]

ngxson and others added 10 commits July 15, 2024 19:23

fix ci (ggerganov#8494)

4db8f60

llama : valign + remove unused ftype (ggerganov#8502)

0efec57

export-lora : handle help argument (ggerganov#8497)

37b12f9

The --help option on export-lora isn't accepted as valid. The help still gets displayed by default, but the script exits with an error message and nonzero status.

make/cmake: add missing force MMQ/cuBLAS for HIP (ggerganov#8515)

5e116e8

CPU/CUDA: Gemma 2 FlashAttention support

cadda27

apply logit_softcap to scale in kernel

fd2539d

disable logit softcapping tests on Metal

c325464

github-actions bot added Nvidia GPU testing examples python server ggml labels Jul 17, 2024

Nexesenex merged commit ebf3686 into Nexesenex:lcpp_pr_flash_gemma Jul 17, 2024
46 of 52 checks passed

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Fattn logit softcap #244

Fattn logit softcap #244

Nexesenex commented Jul 17, 2024

Fattn logit softcap #244

Fattn logit softcap #244

Conversation

Nexesenex commented Jul 17, 2024