Tool MCP: wiki_index_metadata — Adopción de pages solo-repo a D1
Tool MCP: wiki_index_metadata — Adopción de pages solo-repo a D1
Resumen
wiki_index_metadata es una tool MCP introducida en la sesión 18 del supercontexto (commit a575a28, 2026-04-23). Su propósito es adoptar en D1 páginas wiki que ya existen como .md en el repo pero que nunca fueron registradas en bib_wiki_pages — ya sea porque se crearon fuera del flujo MCP (backfill manual, subagentes directos) o porque el registro D1 se perdió.
A diferencia de wiki_create_page, esta tool no escribe ningún archivo al repo. Solo ejecuta un INSERT en la tabla bib_wiki_pages a partir del front-matter ya parseado que el llamador le proporciona.
Contexto: el problema que resuelve
Durante la sesión 17 (backfill de crearack-tech y workspace-*), 7 páginas legacy fueron escritas directamente como .md canónicos en src/content/wiki/ sin pasar por wiki_create_page. Esto las dejó en un estado solo-repo: el .md existe en git pero el registro D1 no, lo que las hace invisibles para bib_search_nodes, bib_ask y cualquier query sobre el grafo.
El flujo de reconcile_wiki_to_d1 en bib_ingest.py ya existía para sincronizar actualizaciones, pero si llamaba a wiki_update_page sobre una página que no estaba en D1, recibía "not found" y simplemente la contabilizaba como skipped_not_in_d1. Con la sesión 18, ese branch de "not found" ahora intenta un upsert via wiki_index_metadata antes de rendirse.
Definición de la tool
Handler: wikiIndexMetadata — functions/api/mcp/handlers/wiki.ts
async function wikiIndexMetadata(
db: D1Database,
args: Record<string, unknown>,
env: Env,
): Promise<string>
Registrado en el router de handleWiki bajo el case 'wiki_index_metadata'.
Input schema (functions/api/mcp/tools.ts)
| Campo | Tipo | Obligatorio | Descripción |
|---|---|---|---|
slug | string | ✅ | Slug kebab de la page. Regex: ^[a-z0-9][a-z0-9-]*[a-z0-9]$. Debe tener .md en src/content/wiki/. |
front_matter | object | ✅ | Campos del front-matter ya parseados. type es obligatorio (validado contra VALID_TYPES). |
front_matter.type | string | ✅ | Tipo de page: entity_page, feature_page, decision_page, concept_page, incident_page, runbook_page. |
front_matter.title | string | — | Título legible. Default: el propio slug. |
front_matter.status | string | — | active o draft. Default: active. |
front_matter.owner | string | — | Default: el valor de actor. |
front_matter.tags | array | — | Lista de tags. Default: []. |
front_matter.sources | array | — | Lista de fuentes. Default: []. |
front_matter.related | array | — | Slugs relacionados. Default: []. |
front_matter.supersedes | string | — | Slug que esta page reemplaza (opcional). |
actor | string | — | Quién adopta la page. Default: agent-ingest-reconcile. |
Output
// Éxito (INSERT realizado):
{ "indexed": true, "slug": "...", "page_id": 123, "content_hash": "abc..." }
// Idempotencia (ya existía en D1):
{ "already_indexed": true, "slug": "...", "page_id": 42 }
// Error de validación:
{ "error": "slug must be lowercase kebab (a-z, 0-9, -)" }
{ "error": "front_matter.type required. Must be one of: ..." }
Comportamiento detallado
1. Validación de slug y tipo
El handler valida el slug con regex y el type contra el set VALID_TYPES antes de cualquier I/O. Devuelve error JSON inmediato si fallan.
2. Idempotencia
SELECT id FROM bib_wiki_pages WHERE slug = ?
Si el slug ya existe en D1, la tool devuelve { already_indexed: true } y no ejecuta ningún INSERT. Esto hace la llamada segura en retries y races.
3. Content hash (best-effort)
Si el Worker tiene GH_PAT disponible en env, intenta leer el .md de GitHub API para calcular el sha256 del contenido real. Si falla (sin token, archivo ausente, timeout), el content_hash queda vacío (''). El INSERT continúa de todas formas.
4. INSERT en D1
INSERT INTO bib_wiki_pages
(slug, type, status, title, owner, file_path, tags, sources, related, supersedes, content_hash)
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
file_path se deriva como src/content/wiki/<slug>.md.
5. Log de auditoría
Tras el INSERT, escribe un registro en bib_wiki_log:
INSERT INTO bib_wiki_log (operation, actor, page_slug, summary, artifacts)
VALUES ('create', ?, ?, 'Indexed existing repo page (metadata-only, no new content)', ?)
6. Registro en LOGGED_TOOLS
La tool queda registrada en functions/api/mcp/index.ts como:
wiki_index_metadata: { action: 'index', entity_type: 'wiki_page' }
Extensión de reconcile_wiki_to_d1 (sesión 18)
Antes (sesión 17 y anterior)
Cuando wiki_update_page retornaba "not found":
if "not found" in err:
stats["skipped_not_in_d1"] += 1
continue
Después (sesión 18)
if "not found" in err:
fm_type = fm.get("type")
if not isinstance(fm_type, str) or not fm_type:
stats["skipped_not_in_d1"] += 1
continue
# Fallback: auto-INSERT via wiki_index_metadata
index_payload = {
"type": fm_type,
"title": fm.get("title") or slug,
"status": fm.get("status") or "active",
"owner": fm.get("owner") or "agent-backfill",
"tags": fm.get("tags") or [],
"related": fm.get("related") or [],
}
raw2 = call_mcp("wiki_index_metadata", {"slug": slug, "front_matter": index_payload, "actor": "agent-ingest-reconcile"}, mcp_token)
result2 = json.loads(raw2) if raw2 else {}
if result2.get("indexed"):
stats["indexed_new"] += 1
stats["pages"].append(slug)
elif result2.get("already_indexed"):
stats["skipped_not_in_d1"] += 1
else:
stats["failed"] += 1
Prerrequisito: el front-matter del .md en el repo debe incluir type. Si falta, se sigue contando como skipped_not_in_d1.
Nuevo contador: indexed_new
El dict stats ahora incluye:
stats = {
"reconciled": 0, # Pages actualizadas en D1 (ya existían)
"indexed_new": 0, # Pages adoptadas de solo-repo a D1 (nuevo sesión 18)
"skipped_not_in_d1": 0,
"failed": 0,
"pages": [],
}
El log summary refleja ambos:
[Reconcile] D1 sync: 3 updated, 2 newly indexed, 1 skipped, 0 failed
Y el skip_summary al MCP:
reconcile: 5 D1 pages (3 updated / 2 new)
Casos de uso
| Caso | Cómo usarlo |
|---|---|
| Pages legacy backfill (7 pages sesión 17) | touch el .md → el próximo push activa reconcile_wiki_to_d1 → fallback automático via wiki_index_metadata |
| Subagente que escribió .md directo | Llamar manualmente wiki_index_metadata(slug, front_matter) desde Claude Desktop |
| Recovery tras borrado accidental del registro D1 | wiki_index_metadata sin tocar el repo |
| Ingest manual de page externa | Con el front-matter mínimo (type obligatorio) |
Invariantes y limitaciones
- No escribe al repo: si el
.mdno existe, el registro D1 apuntará a unfile_pathfantasma. Siempre verificar que el.mdexista antes de llamar. typees obligatorio: sin él, la tool devuelve error yreconcilelo cuenta comoskipped.- Idempotente: llamadas repetidas con el mismo slug son safe (no duplican el registro).
- Content hash opcional: si el Worker no tiene
GH_PAT, el hash queda vacío, lo que puede generar falsas diferencias en reconciles futuros. - No actualiza: si el slug ya existe en D1 con metadata obsoleta, esta tool no la actualiza. Usar
wiki_update_pagepara eso.
Véase también
- [[feature—supercontext—reconcile-d1-repo]] — consumidor principal de la tool; el fallback del reconcile llama a
wiki_index_metadatacuandowiki_update_pagedevuelvenot found - [[concept—biblioteca—supercontexto]] — arquitectura global del sistema de bibliotecarios al que pertenece esta tool
- [[entity—mcp—tool—wiki-lint-bulk]] — tool hermana en el set MCP de wiki que consume la misma tabla
bib_wiki_pages - [[entity—mcp—tool—wiki-utility-recompute]] — tool hermana que recalcula métricas sobre las páginas adoptadas por esta tool