Servicio llama-help: llama.cpp self-host para síntesis RAG (Gemma 4 E2B Q4_0)
Servicio llama-help
⚠️ SERVICIO DADO DE BAJA · s53 (08-05-2026)
Este servicio ya no existe. El servidor
pve-epyc-02fue wipe (NIST SP 800-88) + cancelado en Hetzner Robot. El Cloudflare Tunnelllama-help.crearack.com, el DNS CNAME y el secretLLAMA_HELP_API_KEYfueron eliminados.El backend de síntesis del Help Widget vive ahora en Google AI Studio Paid Tier (
gemma-4-26b-a4b-it). Ver [[decision—20260511—cleanup-openrouter-monoproveedor-google-genai]] y el ADR superseded [[decision—20260507—gemma4-e2b-self-host-synthesis]].Página conservada como referencia histórica.
Servidor llama.cpp self-hosted que ejecutaba Gemma 4 E2B Q4_0 en LXC 100 de pve-epyc-02, expuesto públicamente como https://llama-help.crearack.com vía Cloudflare Tunnel. Fue el backend de síntesis del Help Widget de CreaRack Pro entre el 2026-05-07 y el 2026-05-08.
Propósito
Provee un endpoint OpenAI-compatible (/v1/chat/completions) que synthesizeAnswer en archivo-core.ts llama para generar respuestas RAG a preguntas de usuarios del Help Widget. Reemplaza Google AI Studio (Gemma 4 26B-A4B-IT) como backend de síntesis.
Infraestructura
| Campo | Valor |
|---|---|
| Host | LXC 100 en pve-epyc-02 (Proxmox, EPYC) |
| Puerto interno | 8081 (localhost only) |
| URL pública | https://llama-help.crearack.com/v1/chat/completions |
| Exposición | Cloudflare Tunnel (cloudflared) |
| Systemd unit | llama-server-e2b.service |
| Modelo | Gemma 4 E2B Q4_0 (unsloth/gemma-4-E2B-it-GGUF) |
| Threads | 16 |
| Contexto | 16384 tokens |
| Multimodal | No (sin mmproj, text-only) |
Parámetros de sampling (oficial Gemma 4)
{
"temperature": 1.0,
"top_p": 0.95,
"top_k": 64,
"max_tokens": 1024
}
Características clave
--jinja: separa automáticamente chain-of-thought enreasoning_content. El cliente (archivo-core.ts) solo leechoices[0].message.content— el scratchpad nunca llega al usuario.- API OpenAI-compatible:
messages: [system, user]— no requiere adaptadores especiales. - Autenticación: Bearer token via
LLAMA_HELP_API_KEY(CF Pages secret).
Rendimiento observado
- ~43 tok/s en pve-epyc-02 (EPYC multi-core)
- Latencia típica: 5–15 s (vs ~28 s del modelo 26B en AI Studio)
Integración con el código
Variable de entorno (types.ts)
export interface Env {
// ...
LLAMA_HELP_API_KEY?: string;
// ...
}
Guard en endpoints
// biblioteca/ask.ts y archivo.ts
if (!env.LLAMA_HELP_API_KEY) {
return Response.json({ error: 'LLAMA_HELP_API_KEY not configured — synthesis unavailable' }, { status: 503 });
}
Llamada en archivo-core.ts
const LLAMA_HELP_URL = 'https://llama-help.crearack.com/v1/chat/completions';
export const SYNTHESIS_MODEL = 'gemma-4-E2B-it-Q4_0';
const res = await fetch(LLAMA_HELP_URL, {
method: 'POST',
headers: {
'Content-Type': 'application/json',
Authorization: `Bearer ${apiKey}`,
},
body: JSON.stringify({
messages: [
{ role: 'system', content: systemPrompt },
{ role: 'user', content: userPrompt },
],
max_tokens: 1024,
temperature: 1.0,
top_p: 0.95,
top_k: 64,
}),
});
const answer = data.choices?.[0]?.message?.content || 'Sin respuesta del modelo.';
Prompt del sistema
El systemPrompt enviado instruye al modelo con estilo coloquial específico para CreaRack Pro:
- Tono natural, frases cortas, sin jerga técnica innecesaria.
- Negritas
**...**solo para términos literales de UI (máx 3 por respuesta). - Una sola cita
[N]por párrafo, nunca tras cada frase. - Listas solo si hay >3 pasos.
- Idioma: el de la pregunta.
Disponibilidad y riesgos
- Sin fallback automático: si pve-epyc-02 o el CF Tunnel caen,
bib_askretorna 503. LLAMA_HELP_API_KEYrequerida: debe estar configurada en CF Pages antes de que el endpoint funcione.- Calidad vs 26B: modelo 2B — suficiente para RAG con chunks buenos, puede ser menos preciso en razonamiento complejo.
Historial de cambios
| Fecha | Evento |
|---|---|
| 2026-05-07 | Creado. Reemplaza Google AI Studio como backend de síntesis (commit 8600d88). |
| 2026-05-08 (s53) | Dado de baja: pve-epyc-02 wipe NIST + cancelado en Hetzner Robot. Tunnel llama-help.crearack.com + DNS CNAME + LLAMA_HELP_API_KEY eliminados. Help Widget revertido a Google AI Studio Paid Tier (gemma-4-26b-a4b-it). |
Véase también
- [[decision—20260507—gemma4-e2b-self-host-synthesis]]
- [[entity—biblioteca—handler—archivo-core]]
- [[decision—20260428—gemma-4-via-openrouter-migration]]
- [[feature—biblioteca—bib-ask-synthesis]]
- [[concept—infra—self-host-llm]]