Packages
llama_cpp_ex
0.8.23
0.8.36
0.8.35
0.8.34
0.8.33
0.8.32
0.8.31
0.8.28
0.8.27
0.8.26
0.8.25
0.8.24
0.8.23
0.8.22
0.8.21
0.8.20
0.8.19
0.8.18
0.8.17
0.8.16
0.8.15
0.8.14
0.8.13
0.8.12
0.8.11
0.8.10
0.8.9
0.8.8
0.8.7
0.8.6
0.8.5
0.8.4
0.8.3
0.8.2
0.8.1
0.8.0
0.7.9
0.7.8
0.7.7
0.7.6
0.7.5
0.7.4
0.7.3
0.7.2
0.7.0
0.6.14
0.6.13
0.6.12
0.6.11
0.6.10
0.6.9
0.6.8
0.6.7
0.6.6
0.6.5
0.6.4
0.6.3
0.6.1
0.6.0
0.5.0
0.4.4
0.4.3
0.4.2
0.4.1
0.3.0
0.2.0
Elixir bindings for llama.cpp — run LLMs locally with Metal, CUDA, Vulkan, or CPU acceleration.
Current section
Files
Jump to
Current section
Files
lib/llama_cpp_ex/model_manager/backend.ex
defmodule LlamaCppEx.ModelManager.Backend do
@moduledoc """
Behaviour for the model I/O the manager performs on the write path.
The default implementation is `LlamaCppEx.ModelManager.ModelIO`, which delegates
to `LlamaCppEx.Hub`, `LlamaCppEx.Model`, and `LlamaCppEx.Server`. Tests inject a
fake via the `:io` start option to exercise load/unload lifecycle without real
GGUF files.
Inference dispatch (`generate`/`stream`/`chat`/`embed`) does NOT go through this
behaviour — it reads the ETS table directly from the caller and calls the
relevant module, keeping the manager process off the hot path.
"""
alias LlamaCppEx.Model
alias LlamaCppEx.ModelManager.Entry
@doc """
Resolves a source to a local file path and its byte size, downloading from the
Hub if needed.
"""
@callback resolve_source(Entry.source(), keyword()) ::
{:ok, String.t(), non_neg_integer()} | {:error, term()}
@doc "Loads a model directly (for `:direct` mode)."
@callback load_model(String.t(), keyword()) :: {:ok, Model.t()} | {:error, term()}
@doc "Starts a backing `LlamaCppEx.Server` for `id` (for `:server` mode)."
@callback start_server(id :: term(), path :: String.t(), keyword()) ::
{:ok, pid()} | {:error, term()}
@doc "Stops a backing server, dropping its context and model refs."
@callback stop_server(pid()) :: :ok
end