llama.cpp

Files

T

Sirui He 073bb2c20b mtmd : add MERaLiON-2 multimodal audio support (#21756 )

* mtmd : add MERaLiON-2 multimodal audio support

Adds support for A*STAR's MERaLiON-2 audio-language model (3B and 10B)
to the multimodal framework.

Architecture:
- Whisper large-v2 encoder for audio feature extraction
- Gated MLP adaptor: ln_speech -> frame stack (x15) -> Linear+SiLU -> GLU -> out_proj
- Gemma2 3B / 27B decoder

The mmproj GGUF is generated via convert_hf_to_gguf.py --mmproj on the full
MERaLiON-2 model directory (architecture: MERaLiON2ForConditionalGeneration).
The decoder is converted separately as a standard Gemma2 model after stripping
the text_decoder. weight prefix.

New projector type: PROJECTOR_TYPE_MERALION

Supports tasks: speech transcription (EN/ZH/MS/TA), translation, spoken QA.

Model: https://huggingface.co/MERaLiON/MERaLiON-2-3B
       https://huggingface.co/MERaLiON/MERaLiON-2-10B

* simplify comments in meralion adaptor

* meralion: use format_tensor_name, ascii arrows in comments

2026-04-11 14:15:48 +02:00

batched-bench

common : move up common_init() and fix Windows UTF-8 logs (#21176 )

2026-03-31 12:53:41 +02:00

cli

server: save and clear idle slots on new task (--clear-idle) (#20993 )

2026-04-03 19:02:27 +02:00

completion

server: save and clear idle slots on new task (--clear-idle) (#20993 )

2026-04-03 19:02:27 +02:00

cvector-generator

common : move up common_init() and fix Windows UTF-8 logs (#21176 )