llama-server · runbook técnico
Estado: producción operativa desde sesión 45 (01-05-2026)
Hospedaje: LXC 100 crearack-llama-epyc-02 sobre Proxmox VE 9.1.9 en Hetzner EPYC 7502P
Versión activa: llama.cpp b8984 build from source con -DGGML_NATIVE=ON (commit 45155597a)
Modelo: Gemma 4 26B-A4B-it Q4_K_M + mmproj-F16 (visión)
1. Identidad y endpoints
| Campo | Valor |
|---|---|
| Hostname Tailscale (LXC) | crearack-llama-epyc-02 |
| IP Tailscale (LXC) | 100.80.142.60 |
| Puerto | :8080 |
| URL | http://crearack-llama-epyc-02:8080 o http://100.80.142.60:8080 |
| TLS | No (Tailscale ya cifra extremo a extremo con WireGuard) |
| Auth | Ninguna actualmente. Solo accesible vía tailnet CreaRackSL@ |
| API compatible | OpenAI Chat Completions (/v1/chat/completions) |
Rutas expuestas por llama-server
| Ruta | Tipo | Para qué |
|---|---|---|
GET / | HTML | Frontend web embebido (chat UI) |
GET /health | JSON | Health check ({"status":"ok"}) |
GET /props | JSON | Parámetros default + slots disponibles |
GET /metrics | Prometheus | Métricas runtime (slots, tokens/seg, etc.) |
POST /v1/chat/completions | OpenAI-compatible | Chat completions (con/sin stream, multimodal) |
POST /v1/completions | OpenAI-compatible | Completion legacy |
POST /v1/embeddings | OpenAI-compatible | Embeddings (no usado actualmente) |
POST /completion | Nativo | API legacy llama.cpp |
POST /tokenize | Nativo | Convertir texto a tokens |
POST /detokenize | Nativo | Convertir tokens a texto |
POST /infill | Nativo | Completion in-fill (FIM) |
2. Configuración systemd
/etc/systemd/system/llama-server.service dentro del LXC 100:
[Unit]
Description=llama-server (CreaRack Auto-Plan visión, b8984 native Zen 2)
After=network-online.target
Wants=network-online.target
[Service]
Type=simple
Restart=on-failure
RestartSec=5
Environment="LD_LIBRARY_PATH=/opt/llama.cpp/llama-b8984"
ExecStart=/opt/llama.cpp/llama-b8984/llama-server \
--model /var/lib/llamacpp/models/gemma-4-26B-A4B-it-UD-Q4_K_M.gguf \
--mmproj /var/lib/llamacpp/models/mmproj-F16.gguf \
--host 0.0.0.0 --port 8080 \
--ctx-size 16384 --threads 24 --threads-batch 24 \
--no-mmap --jinja
LimitNOFILE=65536
[Install]
WantedBy=multi-user.target
Parámetros relevantes
| Flag | Valor | Razón |
|---|---|---|
--model | gemma-4-26B-A4B-it-UD-Q4_K_M.gguf (16 GB) | Q4_K_M quantization, equilibrio calidad/RAM |
--mmproj | mmproj-F16.gguf (1.2 GB) | Vision projection, F16 precision |
--ctx-size | 16384 | Contexto máximo (16K tokens). Auto-Plan usa ~1.4K, holgura sobrada |
--threads | 24 | Cores físicos pinneados al LXC (cgroup cpuset.cpus = 8-31) |
--threads-batch | 24 | Threads para prompt processing |
--no-mmap | activo | Carga el modelo entero en RAM (16 GB). Sin esto, mmap’a archivo y pierde performance al primer token |
--jinja | activo | Usa Jinja chat templates (necesario para Gemma 4 con thinking control) |
Ajustes vivientes
# Restart tras cambio de unit
pct exec 100 -- systemctl daemon-reload
pct exec 100 -- systemctl restart llama-server.service
# Status
pct exec 100 -- systemctl status llama-server.service
# Logs en vivo
pct exec 100 -- journalctl -fu llama-server.service
3. Switch entre versiones de llama.cpp
Coexisten dos versiones instaladas en /opt/llama.cpp/:
pct exec 100 -- ls -d /opt/llama.cpp/llama-b*
# /opt/llama.cpp/llama-b8984 ← build NATIVE Zen 2 (activa)
# /opt/llama.cpp/llama-b8994 ← release Ubuntu x64 (haswell autodetect)
/usr/local/bin/llama-switch
Cambia la versión activa modificando el systemd unit + restart:
# Sin args: lista versiones + muestra activa
pct exec 100 -- llama-switch
# Cambiar a otra versión
pct exec 100 -- llama-switch b8994
# Cambio: llama-b8984 → llama-b8994
# OK · activa: llama-b8994
# Rollback rápido (~5 segundos restart)
pct exec 100 -- llama-switch b8984
/usr/local/bin/llama-update-native
Build NATIVE Zen 2 de la última tag estable de llama.cpp:
# Sin args: descarga + build última tag
pct exec 100 -- llama-update-native
# Última tag estable: b9050
# Checkout b9050...
# Build NATIVE Zen 2 (~5-8 min con 24 cores)...
# Instalado: /opt/llama.cpp/llama-b9050
# Para activar: llama-switch b9050
# Versión específica
pct exec 100 -- llama-update-native b9050
El script clona https://github.com/ggml-org/llama.cpp.git a /opt/llama.cpp/source, hace git fetch --tags, checkout, y compila con -DCMAKE_BUILD_TYPE=Release -DGGML_NATIVE=ON -DLLAMA_CURL=ON -DBUILD_SHARED_LIBS=ON. Output a /opt/llama.cpp/llama-<TAG>/. NO activa la nueva versión automáticamente — hay que hacer llama-switch <TAG> después.
Workflow típico de upgrade:
pct exec 100 -- llama-update-native # ~5-8 min build
pct exec 100 -- llama-switch <TAG_NEW> # 5 seg restart
# Validar con plano 70 racks Auto-Plan UI
# Si OK queda; si regresión perf:
pct exec 100 -- llama-switch b8984 # rollback
4. Decisiones técnicas
Build NATIVE Zen 2 vs release binario
Probado empíricamente en sesión 45 con plano 70 racks (5300+ tokens decode):
| Setup | Prompt tps | Decode tps sustained | Tiempo total |
|---|---|---|---|
b8984 build NATIVE -DGGML_NATIVE=ON | 91.85 | ~26-27 | ~3:30 |
b8994 release con libggml-cpu-haswell.so autodetect | 90.12 | ~18-19 | ~5:00 |
Diferencia decode sustained: ~30 % peor con haswell vs NATIVE en este modelo+hardware. Razones probables: build NATIVE usa instrucciones específicas Zen 2 (AES-NI, SHA-NI, REPACK) además de las comunes con Haswell (AVX2, FMA, F16C); en outputs largos donde el matmul domina, esas extras suman.
Decisión vinculante: para este host, siempre build NATIVE vía llama-update-native. NO usar release pre-compilado salvo emergencia. Si se hace, documentar la regresión esperada.
—threads 24 (no 32)
LXC tiene cgroup pinning cpuset.cpus = 8-31 (24 cores físicos). El host reserva cores 0-7 (8 cores) para Proxmox kernel + I/O + network — necesario para no saturar el host bajo carga llama-server al 100 %.
--threads 24 matchea exactamente los cores disponibles. Subir a más threads (32) provocaría context switching innecesario sin ganar throughput porque el bottleneck es bandwidth de RAM, no cores.
—no-mmap
Sin esta flag, llama-server hace mmap del GGUF y los primeros tokens son MUY lentos (page faults forzando lecturas desde disco). Con --no-mmap carga los 16 GB enteros en RAM al arrancar (~30 segundos) y a partir de ahí cada request es full speed.
5. Observabilidad
Logs
# Stream live
pct exec 100 -- journalctl -fu llama-server.service
# Filtrar requests + timing
pct exec 100 -- journalctl -fu llama-server.service \
--since now -o cat \
| grep -E "POST /v1|prompt eval time|predicted_per_second|done request"
Métricas Prometheus
curl http://crearack-llama-epyc-02:8080/metrics
Incluye llamacpp:n_decode_tokens_total, llamacpp:prompt_tokens_total, llamacpp:tokens_predicted_per_second, etc. Pendiente integrar al stack de monitoring CreaRack PROD (VictoriaMetrics) — futuro.
Health check externo
curl -s http://crearack-llama-epyc-02:8080/health
# {"status":"ok"}
Útil para incluir en UptimeRobot interno o sensor del workspace.
6. Auto-Plan PROD — integración
CreaRack PROD apunta aquí vía Dokploy env vars:
AUTOPLAN_PROVIDER=ollama
OLLAMA_BASE_URL=http://100.80.142.60:8080
OLLAMA_MODEL=gemma-4-26B-A4B-it-UD-Q4_K_M.gguf
EDGE_AI_PROVIDER=openrouter
AUTOPLAN_PROVIDER=ollamaes histórico — el driver Python encore/services/ai_providers/se llama “ollama” pero apunta a:8080llama-server (compatible OpenAI). NO hay Ollama instalado en este server (purgado en sesión 44).
Cambio de versión llama.cpp NO requiere tocar Dokploy. La env var apunta al puerto + IP, no al binario.
Para reapuntar PROD a otro server (rollback de hardware o test):
- Panel Dokploy →
crearack-pro→ Environment. - Cambiar
OLLAMA_BASE_URLa la IP Tailscale del nuevo destino. - Save → Restart containers.
NO modificar
compose.ymldirectamente — env vars en panel Dokploy (footgun documentado en setup PROD).
7. Modelos GGUF instalados
/var/lib/llamacpp/models/ (bind-mount al subdataset ZFS llamacpp/models en host):
| Archivo | Tamaño | Uso |
|---|---|---|
gemma-4-26B-A4B-it-UD-Q4_K_M.gguf | 16 GB | producción Auto-Plan + chat |
mmproj-F16.gguf | 1.2 GB | visión asociado al 26B |
E4B/gemma-4-E4B-it-Q6_K.gguf | 6.6 GB | descartado calidad 20 % |
E4B/mmproj-F16.gguf | 0.92 GB | visión asociado al E4B |
Pendiente s46+: probar Q6_K (21 GB) y Q8_0 (25 GB) del 26B-A4B para subir 70→80 % calidad. Cabe sin problema en 64 GB del LXC. Descargar de unsloth/gemma-4-26B-A4B-it-GGUF en HuggingFace.
Para añadir un nuevo modelo:
pct exec 100 -- bash -c '
cd /var/lib/llamacpp/models
wget https://huggingface.co/<repo>/<archivo>.gguf
'
# Editar systemd unit para apuntar al nuevo .gguf
# pct exec 100 -- systemctl restart llama-server.service
Para hot-swap de modelo sin reiniciar (test A/B), arrancar segunda instancia en otro puerto:
# Variante: copiar systemd unit a llama-server-test.service apuntando al modelo alternativo + puerto 8081
# Permite comparar simultáneamente sin tocar producción
8. Troubleshooting
Servicio no arranca tras restart
pct exec 100 -- journalctl -u llama-server.service --no-pager -n 30
Causas habituales:
- Modelo no encontrado:
gguf_init_from_file: failed to open GGUF file (No such file or directory). Verificar bind-mount LXCmp0: /llamacpp/models,mp=/var/lib/llamacpp/modelsy que/llamacpp/models/<gguf>existe en el host. - Memoria insuficiente: el LXC tiene 64 GB. Modelo Q4_K_M usa ~17-18 GB en runtime. Si compartes con otros procesos pesados, OOM. Comprobar
free -hen el LXC. - Permisos: si ZFS dataset no tiene chown 100000 (uid mapping unprivileged), LXC no puede leer.
chown 100000:100000 /llamacpp/modelsen el host.
Auto-Plan PROD no llega al llama-server
ssh root@crearack.com "docker exec crearack-pro-zcmvsl-web-1 \
curl -s -m 5 http://100.80.142.60:8080/health"
Si timeout: el contenedor PROD ha perdido Tailscale o está fuera del tailnet. Verificar tailscale status en el host PROD.
Performance regresión sin cambio aparente
Posibles causas:
- Otra LXC/VM acaparando cores 8-31. Verificar con
htopen host. - Memoria swap activada (debería estar a 0 — verificar
free -h). - CPU governor cambió a
powersave. Verificarcat /sys/devices/system/cpu/cpu0/cpufreq/scaling_governor(debe decirperformance). - Modelo cargado en disco en vez de RAM (si quitaron
--no-mmappor error).
9. Capas de seguridad
- Internet público: 0 acceso. UFW del host con default deny incoming, solo
:22 on tailscale0. LXC sin IP pública. - Tailnet
CreaRackSL@: cualquier device logueado puede llegar a:8080sin auth. Devices actuales:yogaedu,pixel-9-pro-xl,pve-epyc-02,crearack-llama-epyc-02,crearack-prod. - API key: NO activada. Para activar, añadir
--api-key <token>al systemd unit + actualizar Dokploy PROD para enviarAuthorization: Bearer <token>. - Tailscale ACL: por defecto permit all entre nodos del mismo tailnet. Para restringir cuáles devices pueden llegar al
:8080, editar policy file enhttps://login.tailscale.com/admin/acls.
Decisión actual (s45): se mantiene sin API key porque el tailnet es controlado (2-3 personas, pocos devices). Re-evaluar si entran becarios o más devs.
10. Referencias
- llama.cpp upstream: https://github.com/ggml-org/llama.cpp
- Releases: https://github.com/ggml-org/llama.cpp/releases
- Server docs: https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md
- Modelos GGUF Gemma 4: https://huggingface.co/unsloth/gemma-4-26B-A4B-it-GGUF
- Server hospedaje: [[crearack-tech—admin—proxmox-epyc-setup]]
- Guía usuario: [[workspace—guias—llama-acceso-uso]]
Véase también
- [[crearack-tech—admin—proxmox-epyc-setup]] — host Proxmox + ZFS + LXC (madre infra)
- [[crearack-tech—admin—server-management]] — gestión Hetzner general
- [[crearack-tech—admin—monitoring-tools]] — herramientas observabilidad CreaRack
- [[workspace—guias—llama-acceso-uso]] — guía usuario coloquial del chat
- [[workspace—agentes-ia]] — comparativa agentes IA del equipo (Claude, ChatGPT, Gemini, este Llama)