server : speculative checkpointing (#19493)

* server : speculative decoding using checkpoints

* server : fix draft check with checkpoints

* server : rename spec vars

* server : log levels

* server : refactored spec logic to speculative.cpp

* server : renamed spec checkpoints option

* server : fix spec checkpoints, logging

* speculative : checkpoints with draft model, logging

* server : n_tokens_cur and create_checkpoint in draft

* server : fix server_speculative_callback (slot.id)

* spec : fix ngram-map/begin idx_last_check

* spec : init ckpt (begin() wasn't called)

* chore: update webui build output

* server : restore sampler in spec checkpoint and clear mem

* cont : avoid --spec-use-checkpoints argument

* cont : remove server_prompt_checkpoint_with_size

* spec : rename (leave_draft_state)

* cont : clean-up

* cont : do not ignore partial drafts even if the are short

* cont : spec callback owned by session

* cont : simplify

* cont : avoid empty speculative session

* cont : simplify

* cont : simplify

* cont : enable mtmd speculative decoding

* cont : keep the spec sampler alive

* cont : simplify

* cont : fix nullptr deref + draft checkpoints

* cont : remove common_speculative_accept_response

* cont : remove callback

* cont : simplify

* cont : minor

* cont : simplify

* cont : fix accepted number

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

This commit is contained in:

Sascha Rogmann

2026-04-19 09:24:06 +02:00

committed by

GitHub

parent 91fef95362

commit 455d8e4be8

10 changed files with 421 additions and 180 deletions

									
										tools/server/server-common.h
									
		+3
		-1
	
												View File
												
				@@ -190,7 +190,9 @@ public:

				    void insert(const llama_tokens & inp_tokens);

				    // for compatibility with speculative decoding, ctx shift, slot save/load

				    const llama_tokens & get_text_tokens() const;

				    const llama_tokens & get_tokens() const;

				    llama_tokens get_text_tokens() const;

				    // for compatibility with speculative decoding

				    void set_token(llama_pos pos, llama_token id);