Files
llm/docs/quality-audit-logging-plan.md
T

9.8 KiB

Quality And Audit Logging Plan

Goal

Build a minimal but complete logging and telemetry contour that lets us:

  • reconstruct any user turn end to end;
  • compare model, prompt, RAG, and tool variants;
  • detect regressions and failure clusters;
  • export clean datasets for evals and later fine-tuning;
  • do this without leaking secrets into long-term storage.

This plan treats logging as a cross-cutting platform capability under core, while each plugin keeps its domain-specific fields and labels.

Current State

What already exists:

  • plugins/1c/agent stores chat messages and payload JSON in SQLite.
  • 1c-agent persists outbound model context, inbound provider response, latency_ms, tool results, and RAG context.
  • trace_id exists at the API level for the agent.

What is missing:

  • no centralized event schema across plugins and services;
  • no guaranteed trace propagation across agent -> adapter/MCP -> model backend;
  • no normalized turn-level analytics tables;
  • no token/cost accounting;
  • no quality labels or review workflow;
  • no retention/sanitization/export policy for long-term analysis.

Target Outcome

For every important request we should be able to answer:

  1. What the user asked.
  2. Which prompt, model, route, tools, and knowledge sources were used.
  3. What the system returned.
  4. Whether the result was successful, partial, wrong, unsafe, or abandoned.
  5. Which change caused quality to improve or regress.

Logging Layers

1. Access Log

Capture every inbound HTTP request and outbound response for platform services.

Required fields:

  • timestamp
  • service
  • instance
  • environment
  • request_id
  • trace_id
  • method
  • path_template
  • status_code
  • duration_ms
  • request_size_bytes
  • response_size_bytes
  • client_type
  • error_code

Notes:

  • Do not store raw request/response bodies here by default.
  • This is for traffic, latency, and error-rate analysis.

2. Turn Audit Log

Capture one normalized record per user-visible turn.

Required fields:

  • turn_id
  • timestamp
  • project_id
  • chat_id
  • plugin
  • service
  • trace_id
  • request_id
  • user_message_id
  • assistant_message_id
  • user_text
  • assistant_text
  • outcome
  • failure_type
  • duration_ms
  • human_review_status

Recommended enums:

  • outcome: success, partial, failure, refused, abandoned
  • failure_type: none, routing, tool_misuse, hallucination, format_error, timeout, provider_error, adapter_error, rag_miss, policy_error, unknown

3. Model Call Log

Capture each actual provider/model request.

Required fields:

  • model_call_id
  • turn_id
  • trace_id
  • provider_id
  • provider_type
  • base_url
  • model_registry_id
  • served_model_name
  • route_name
  • temperature
  • max_tokens
  • prompt_messages_json
  • response_json
  • prompt_tokens
  • completion_tokens
  • total_tokens
  • cost_estimate
  • latency_ms
  • finish_reason
  • cache_hit

Notes:

  • prompt_messages_json is the exact payload sent to the model after prompt assembly and tool/RAG injection.
  • response_json is the raw provider response after secret stripping.

4. Tool Call Log

Capture every tool or adapter call.

Required fields:

  • tool_call_id
  • turn_id
  • trace_id
  • tool_family
  • tool_name
  • target_service
  • request_json
  • response_json
  • status
  • duration_ms
  • retry_count

Recommended tool families:

  • adapter
  • mcp
  • rag
  • internal
  • external-http

5. Retrieval Log

Capture what RAG actually did.

Required fields:

  • retrieval_id
  • turn_id
  • trace_id
  • plugin
  • profile
  • index_version
  • query_text
  • top_k
  • sources_json
  • context_chars
  • retrieval_latency_ms

6. Review Log

Capture human or automated quality judgments.

Required fields:

  • review_id
  • turn_id
  • reviewer
  • review_source
  • score_helpfulness
  • score_correctness
  • score_tool_use
  • score_safety
  • label_primary_issue
  • notes
  • created_at

Recommended review_source values:

  • human
  • eval
  • heuristic
  • llm-judge

Trace Propagation Rules

Every inbound request creates or adopts:

  • request_id
  • trace_id

Propagation rules:

  • agent must pass trace_id and request_id to adapter/MCP/model wrappers;
  • adapter and MCP must echo them in responses and logs;
  • background jobs must generate child spans but preserve the parent trace_id;
  • exported reports must include source trace_id where possible.

If a downstream protocol cannot carry headers directly, include these IDs in the JSON payload envelope.

Secret And Privacy Rules

Never persist raw secrets in long-term logs.

Must redact or hash before storage:

  • API keys
  • bearer tokens
  • cookies
  • passwords
  • connection strings with credentials
  • user-uploaded files containing secrets or personal data

Recommended approach:

  • keep raw payload only in short-lived memory for request execution;
  • write sanitized JSON to persistent audit storage;
  • store a redaction_applied=true flag and optional redaction_rules_version.

Retention Policy

Use three storage horizons.

Hot

  • Purpose: active debugging and operator support.
  • Storage: local SQLite or service-local JSONL.
  • Retention: 7-14 days.

Warm

  • Purpose: quality analytics and regression analysis.
  • Storage: central SQL tables or partitioned JSONL/Parquet under reports/.
  • Retention: 30-90 days.

Cold

  • Purpose: curated datasets and incident forensics.
  • Storage: exported reviewed samples only.
  • Retention: explicit/manual.

Minimal Schema Proposal

Add a new cross-cutting audit storage layer under core.

Suggested logical entities:

  • access_events
  • turn_audit
  • model_calls
  • tool_calls
  • retrieval_events
  • quality_reviews

Suggested implementation path:

  • phase 1: SQLite in each service plus nightly export to JSONL
  • phase 2: normalized central SQLite/Postgres
  • phase 3: dashboards and automated quality reports

Implementation Order

Phase 1. Fast Wins

Goal: get useful forensic visibility with minimal code churn.

Tasks:

  1. Add access logging middleware/pattern to 1c-agent and other HTTP services.
  2. Split current message journal into explicit turn_audit and model_call records in addition to existing message history.
  3. Add stable turn_id and model_call_id.
  4. Start persisting token usage if the provider returns it.
  5. Add redaction helper for secrets before writing payload JSON.
  6. Export daily JSONL snapshots into reports/observability/.

Definition of done:

  • any turn can be reconstructed from one turn_id;
  • any provider failure can be grouped by model, route, and error code;
  • secrets are not written as plain text.

Phase 2. Cross-Service Correlation

Goal: see one trace across agent, adapter, MCP, and model route.

Tasks:

  1. Propagate trace_id and request_id through agent -> adapter/MCP.
  2. Add the same IDs to adapter logs and tool results.
  3. Normalize tool-call audit records.
  4. Record RAG index version and exact retrieved sources.
  5. Add per-turn outcome and failure_type.

Definition of done:

  • a bad answer can be traced to routing, retrieval, tool use, or model output;
  • one trace can be joined across at least agent and adapter/MCP logs.

Phase 3. Quality Loop

Goal: turn logs into a real quality improvement loop.

Tasks:

  1. Add review/export scripts for bad turns and representative samples.
  2. Add quality labels and scoring workflow.
  3. Produce weekly aggregate reports:
    • error rate by plugin
    • failure types by route/model
    • tool success rate
    • latency percentiles
    • token usage and cost
  4. Feed reviewed samples into eval datasets and training candidates.

Definition of done:

  • we can say which model/prompt/route is better using saved evidence;
  • we can build eval slices from production-like failures.

Metrics To Track

Minimum operational metrics:

  • request count
  • non-2xx rate
  • timeout rate
  • p50/p95/p99 latency
  • tool-call success rate
  • empty-response rate
  • average prompt/completion tokens

Minimum quality metrics:

  • turn success rate
  • partial-answer rate
  • hallucination rate
  • tool-misuse rate
  • retrieval-miss rate
  • human review score averages
  • regression rate after route/prompt/model changes

Suggested Repo Additions

Recommended folders:

  • core/observability/
  • docs/runbooks/observability.md
  • scripts/export_audit_logs.py
  • scripts/report_quality_metrics.py

Recommended first shared modules:

  • core/observability/schema.py
  • core/observability/redaction.py
  • core/observability/trace.py
  • core/observability/store.py

1C-Specific Additions

For the 1c plugin, also record:

  • base_id
  • adapter method list
  • saved-state vs active/effective mode
  • resolved object/module selectors
  • truncation/partial flags from adapter
  • extension/base layer hints

This is important because many 1C failures are not generic model failures. They come from incomplete evidence, wrong search strategy, or extension-layer blind spots.

Immediate Next Step

The best next implementation step is:

  1. introduce core/observability with redaction, trace helpers, and JSONL writers;
  2. wire 1c-agent to emit access_events, turn_audit, and model_calls;
  3. add one export script and one weekly report script;
  4. then extend the same contract to other plugins.

This gives us a usable quality loop without waiting for a full observability platform rollout.

Current Implemented Foundation

Already present in the repo:

  • core/observability/trace.py
  • core/observability/redaction.py
  • core/observability/store.py
  • plugins/1c/agent/agent_server.py emits access_events, turn_audit, model_calls, tool_calls, and retrieval_events
  • scripts/summarize_observability_reports.py
  • scripts/report_failed_turns.py
  • scripts/report_quality_metrics.py
  • scripts/report_repeated_failures.py
  • scripts/export_quality_snapshot.py