PROJECT CONTEXT FOR AI Project name: Local LLM Platform Purpose: We are building a local AI platform for running and operating LLM-based services on our own infrastructure, with a strong focus on 1C support. The platform is not just a chat wrapper around models. It is intended to become a reusable engineering base for: - model registry and model selection; - local inference services; - GPU deployment and service switching; - plugin-specific task pipelines; - evaluation and smoke testing; - safe 1C analysis, RAG, and eventually controlled code/change assistance. Main idea: The repository follows a "core + plugins" architecture. - core = shared platform capabilities; - plugins = task-specific domains that can later become standalone services. This is intentionally a middle ground between a monolith and microservices: - today we move faster in one repository; - tomorrow heavy or mature domains can be extracted into separate services. ================================================== 1. WHAT WE ARE BUILDING ================================================== We are building a local multi-plugin LLM platform for these domains: - text; - translation; - audio; - video; - image; - 1C. The long-term goal is: - one shared platform for model lifecycle and deployment; - multiple domain plugins with their own prompts, datasets, evals, adapters, and APIs; - safe operational workflows around real business systems, especially 1C. The repository is not centered on cloud APIs. It is centered on self-hosted/local models and reproducible GPU deployment. Primary GPU deployment target: - docker-gpu.cin.su Shared test Docker host: - docker-test.cin.su Important operational assumption: - heavy model services are not expected to run all at once; - the current GPU workflow is closer to "single heavy active model/service profile" than to "everything always on". ================================================== 2. ARCHITECTURE ================================================== Top-level structure: - core/ - plugins/ - registry/ - scripts/ - docs/ - tests/ - config/ and configs/ - reports/ Meaning of the main parts: core/ - shared, reusable platform logic; - should not depend on plugin-specific business logic. Expected responsibilities in core: - registry; - inference; - training; - evals; - deployment; - storage; - monitoring. plugins/ - domain-specific logic; - each plugin is designed as a future standalone service boundary. registry/ - model cards and templates; - stores metadata about models, adapters, versions, storage paths, resource requirements, and statuses; - does not store model binaries in git. scripts/ - the operational center of the repo; - contains validation, smoke tests, deployment scripts, reporting, indexing, model download helpers, 1C tooling, and service control utilities. docs/ - runbooks, architecture notes, roadmap, API contracts, and research notes. Current architectural principle: - plugins may depend on core; - core must not depend on plugins. ================================================== 3. PLUGINS OVERVIEW ================================================== Text plugin: - general text/chat/code-style use cases; - currently supported by local model registry and vLLM deployment patterns. Translation plugin: - local translation service route exists; - transformers-based service deployment is prepared. Audio plugin: - speech-related plugin; - service route and deployment assets exist; - smoke and readiness tooling exist. Video plugin: - vision/video analysis direction; - service route and deployment assets exist; - intended for frame/video understanding tasks. Image plugin: - generation and editing; - SDXL is the practical current image route; - Qwen image edit experimentation exists, but is much heavier/slower in practice. 1C plugin: - the most mature and strategically important plugin in the repository; - includes RAG, metadata parsing, BSL/module/form analysis, adapter contracts, connector policies, an agent service, training artifacts, evals, and many safety checks. ================================================== 4. THE 1C DIRECTION: WHY IT MATTERS ================================================== The 1C plugin is the deepest part of the project. It is not a simple prompt layer. It is evolving into a safe assistant stack for 1C development and analysis. What the 1C plugin is meant to do: - answer 1C and BSL questions; - help analyze metadata and object structure; - help with read-only 1C queries; - support RAG over 1C documentation and internal knowledge; - inspect forms, modules, templates, and related artifacts; - support controlled change planning; - eventually support fine-tuned adapters/LoRA after enough high-quality examples are collected. Important strategic rule: - first RAG and tooling; - fine-tuning later. This is a deliberate choice. The project is trying to avoid premature fine-tuning before having enough verified, safe, domain-correct examples. ================================================== 5. WHAT HAS ALREADY BEEN IMPLEMENTED ================================================== This repository is already beyond the "empty skeleton" stage. It has real operational substance. Implemented foundation: - architecture and repository layout; - model registry with model cards; - plugin structure for all target domains; - deployment assets for GPU-hosted inference services; - local model chat UI and service control patterns; - large collection of validation and smoke scripts; - runbooks for deployment and operation. Implemented model/platform side: - model registry in registry/model-cards; - plugin model bundle in plugins/model-bundle.yaml; - vLLM deployment assets; - llama.cpp deployment assets; - transformers-based deployment assets for translation/audio/video/image; - runtime profile and GPU profile configuration; - scripts for model download, validation, indexing, status collection, and service management. Implemented UX/operations side: - model chat server and UI for manual model checks; - management console web assets; - platform status collection/reporting; - image generation/edit job flows and report artifacts; - service control endpoints for starting/stopping model-serving stacks. Implemented 1C side: - 1C metadata parsing modules; - XML/form/payload/dbnames/config-related parsing utilities; - 1C RAG corpus preparation and lexical/vector index scripts; - 1C prompts and RAG profile routing; - 1C connector and adapter service boundaries; - 1C MCP adapter code; - 1C agent server with its own web UI; - policy files for read-only queries and change workflow; - schema files for metadata snapshots, BSL snapshots, and moxel registry; - training example scaffolding and LoRA config scaffolding; - many focused checks, smoke tests, analysis scripts, and verification scripts. Evidence of maturity: - about 267 files in scripts/; - about 20 test files under tests/1c alone; - extensive runbook documentation; - many contract-first docs for 1C behavior and write safety. ================================================== 6. CURRENT STATUS AS OF THE DOCUMENTED ROADMAP ================================================== The roadmap file contains a concrete status snapshot dated 2026-06-20. Treat this as the latest documented status inside the repo unless newer evidence is added elsewhere. Documented operating status at that point: - main UI available on LAN; - current GPU profile was "image"; - model chat UI container was running; - image and translation services were online; - text, audio, video, and some 1C-heavy routes were prepared but not always running at the same time; - the platform intentionally used a single-heavy-model approach on RTX 4090 to avoid VRAM conflicts. Documented plugin state: - Text: ready to start via vLLM route. - 1C: ready to start, with a strong GGUF route preferred for manual checks. - Translation: online. - Audio: ready to start. - Video: ready to start. - Image: online, SDXL verified. Documented completed priorities in roadmap: - first-class launcher for Qwen3-Coder Q6 on GPU; - 1C route preference for the strongest practical GGUF model for manual checks; - smoke test additions for more profiles; - image model mode switching; - status report generation from health endpoints. Practical interpretation: - the platform already has operational deployment logic; - however, not every route is meant to be hot simultaneously; - the operator workflow and profile switching are part of the design, not a temporary bug. ================================================== 7. MODEL STRATEGY ================================================== The repository uses a curated "one practical model per plugin" strategy rather than trying to host every possible large model. Examples from the bundle: - text: Qwen3-4B-Instruct; - translation: LMT-60-4B; - audio: Whisper Large V3 Turbo; - video: Qwen2.5-VL 7B; - image: SDXL Base plus SDXL Inpainting; - 1C: Qwen3-Coder GGUF variant as a practical code/1C route. Why this matters: - the project is optimizing for practical local deployment; - VRAM and runtime footprint are first-class constraints; - model selection is tied to plugin purpose, deployment realism, and operator workflow. ================================================== 8. 1C SAFETY MODEL ================================================== This is one of the most important parts for any AI reading this project. The 1C direction is safety-first. The model may: - inspect metadata; - search/read modules; - validate read-only queries; - propose changes; - prepare plans and evidence; - help build reviewable artifacts. The model must not: - directly change a live 1C database; - run destructive queries; - invent metadata when the adapter/connector does not confirm it; - treat effective runtime view as a directly writable surface. Core safety principles already documented in the repo: - read-only first; - full semantic 1C paths are preferred over storage internals; - effective read view and editor provenance must stay separate; - every write must begin with a plan; - direct active configuration writes are forbidden; - saved-state and extension-aware workflows are preferred; - concrete write targets, guards, and validation are mandatory. This means the 1C system is being designed more like a controlled engineering assistant than a free-form coding bot. ================================================== 9. WHAT IS IMPORTANT ABOUT THE 1C ADAPTER/CONNECTOR DESIGN ================================================== The repo already contains a detailed contract for how 1C interaction should work. High-level design ideas: - the adapter works through SQL storage and controlled analysis layers; - XML exports and Form.xml are useful for analysis and learning, but are not the live write transport; - agent-facing APIs should speak in full 1C semantic paths, not raw SQL/CAS/internal offsets; - metadata path resolution and BSL symbol resolution are separate problems; - the adapter must preserve origin/provenance information so we always know whether something belongs to base config, extension, or saved-state. Planned/implemented adapter-side capabilities include: - object resolution; - fact resolution; - BSL symbol resolution; - object context retrieval; - modules/form/template inspection; - write planning; - saved-state targeting and verification; - evidence-preserving decode behavior; - smoke checks and contract checks. This is already a serious design effort, not just a TODO note. ================================================== 10. WHAT REMAINS INCOMPLETE OR STILL EVOLVING ================================================== Even though a lot exists, the project is still in an active build-out phase. Areas that are clearly still evolving: - final extraction of heavy plugins into standalone services; - broader production-grade monitoring and lifecycle management; - full maturity of video/audio/image routes under routine operation; - richer model cabinet / ingest / verification workflows; - more complete 1C live integration loops; - broader eval coverage across plugins; - final training/fine-tuning loop for 1C after verified datasets mature; - more end-to-end workflows that turn read/analysis into safely validated change pipelines. The 1C write path especially is intentionally conservative: - planning and verification are much further along than unrestricted applying; - the project prefers blocked/refused actions over unsafe hidden behavior; - saved-state and extension workflows are being formalized step by step. ================================================== 11. WHAT WE HAVE DONE VS WHAT WE STILL NEED TO DO ================================================== Already done: - defined the platform direction and architecture; - created the plugin structure; - built the model registry concept and model-card inventory; - prepared deployment stacks for local/GPU model serving; - built a manual model-chat/testing UI; - created many operational scripts and checks; - established the 1C plugin as the main advanced domain; - implemented large portions of 1C parsing, RAG, adapter contracts, and safety policies; - documented many workflows and contracts in runbooks and design docs. Still need to do: - continue hardening the platform from "engineering bench" into "repeatable service platform"; - improve standalone service boundaries for heavy plugins; - keep refining GPU profile switching and operational ergonomics; - expand eval coverage and regression checks for all plugins; - continue building the 1C live integration loop with strong provenance and write safety; - improve extension-aware and saved-state-aware write planning and validation; - grow verified 1C datasets; - train and compare 1C-specific LoRA/adapters only after enough trustworthy examples exist; - keep all dangerous operations gated by explicit planning, validation, and rollback-capable flows. ================================================== 12. RECOMMENDED MENTAL MODEL FOR ANOTHER AI ================================================== If you are another AI reading this file, understand the project like this: This repository is a local AI platform with a strong operations and tooling focus. It is not only about prompting models. It is about: - selecting the right local models; - deploying them on constrained GPU hardware; - routing use cases through plugin-specific services; - verifying quality with smoke/eval checks; - building safe interfaces to real-world systems. The 1C area is the flagship domain. Its current priority is safe read/analysis/RAG/tooling with controlled change planning. Unsafe convenience is not acceptable there. You should assume: - reproducibility matters; - deployment realism matters; - model size/VRAM tradeoffs matter; - contracts and runbooks matter; - safety gates matter more than agent autonomy in 1C live workflows. ================================================== 13. HOW TO CONTRIBUTE CORRECTLY ================================================== When continuing work in this repo, prefer these behaviors: - preserve the core + plugins boundary; - keep model binaries and large datasets out of git; - add or update model cards instead of hardcoding assumptions; - prefer operational scripts and reproducible runbooks over one-off manual steps; - for 1C, prefer read-only analysis and explicit planning before any write path work; - preserve provenance/origin information; - use full semantic 1C paths where possible; - do not collapse safe abstractions into raw storage details in agent-facing flows; - add checks/tests/docs together with new behavior; - treat the scripts/ folder as part of the product, not as disposable glue. ================================================== 14. MOST IMPORTANT NEXT STEPS ================================================== The most sensible next steps, based on the repository state, are: 1. Continue strengthening the 1C safe operational loop. - Better write planning. - Better target resolution. - Better saved-state verification. - Better extension provenance and conflict detection. 2. Improve end-to-end readiness of non-1C plugins. - Make text/audio/video/image routes easier to operate repeatedly. - Expand smoke tests and health reporting. - Reduce ambiguity in runtime profiles and active service state. 3. Keep the model registry and deployment inventory clean. - Ensure model cards, bundles, and runtime profiles stay aligned. - Keep practical model choices explicit. 4. Build confidence through validation. - Prefer contract checks, smoke checks, and generated reports. - Keep adding regression coverage where workflows are safety-sensitive. 5. Delay aggressive fine-tuning until the data is worthy. - RAG and tooling first. - Verified examples next. - LoRA/adapters after that. ================================================== 15. SHORT VERSION ================================================== We are building a self-hosted local LLM platform with a plugin architecture. The platform supports text, translation, audio, video, image, and especially 1C. What is already done: - architecture; - model registry; - deployment stacks; - chat/testing UI; - large operations script layer; - strong 1C parsing/RAG/adapter/safety foundation. What is most important now: - mature the platform operationally; - keep non-1C plugins practical to run; - continue deepening the 1C assistant with strict safety, provenance, planning, and validation; - only move to 1C fine-tuning after enough verified domain data exists. End of file.