456 lines
17 KiB
Plaintext
456 lines
17 KiB
Plaintext
PROJECT CONTEXT FOR AI
|
|
|
|
Project name:
|
|
Local LLM Platform
|
|
|
|
Purpose:
|
|
We are building a local AI platform for running and operating LLM-based services on our own infrastructure, with a strong focus on 1C support. The platform is not just a chat wrapper around models. It is intended to become a reusable engineering base for:
|
|
- model registry and model selection;
|
|
- local inference services;
|
|
- GPU deployment and service switching;
|
|
- plugin-specific task pipelines;
|
|
- evaluation and smoke testing;
|
|
- safe 1C analysis, RAG, and eventually controlled code/change assistance.
|
|
|
|
Main idea:
|
|
The repository follows a "core + plugins" architecture.
|
|
- core = shared platform capabilities;
|
|
- plugins = task-specific domains that can later become standalone services.
|
|
|
|
This is intentionally a middle ground between a monolith and microservices:
|
|
- today we move faster in one repository;
|
|
- tomorrow heavy or mature domains can be extracted into separate services.
|
|
|
|
==================================================
|
|
1. WHAT WE ARE BUILDING
|
|
==================================================
|
|
|
|
We are building a local multi-plugin LLM platform for these domains:
|
|
- text;
|
|
- translation;
|
|
- audio;
|
|
- video;
|
|
- image;
|
|
- 1C.
|
|
|
|
The long-term goal is:
|
|
- one shared platform for model lifecycle and deployment;
|
|
- multiple domain plugins with their own prompts, datasets, evals, adapters, and APIs;
|
|
- safe operational workflows around real business systems, especially 1C.
|
|
|
|
The repository is not centered on cloud APIs. It is centered on self-hosted/local models and reproducible GPU deployment.
|
|
|
|
Primary GPU deployment target:
|
|
- docker-gpu.cin.su
|
|
|
|
Shared test Docker host:
|
|
- docker-test.cin.su
|
|
|
|
Important operational assumption:
|
|
- heavy model services are not expected to run all at once;
|
|
- the current GPU workflow is closer to "single heavy active model/service profile" than to "everything always on".
|
|
|
|
==================================================
|
|
2. ARCHITECTURE
|
|
==================================================
|
|
|
|
Top-level structure:
|
|
- core/
|
|
- plugins/
|
|
- registry/
|
|
- scripts/
|
|
- docs/
|
|
- tests/
|
|
- config/ and configs/
|
|
- reports/
|
|
|
|
Meaning of the main parts:
|
|
|
|
core/
|
|
- shared, reusable platform logic;
|
|
- should not depend on plugin-specific business logic.
|
|
|
|
Expected responsibilities in core:
|
|
- registry;
|
|
- inference;
|
|
- training;
|
|
- evals;
|
|
- deployment;
|
|
- storage;
|
|
- monitoring.
|
|
|
|
plugins/
|
|
- domain-specific logic;
|
|
- each plugin is designed as a future standalone service boundary.
|
|
|
|
registry/
|
|
- model cards and templates;
|
|
- stores metadata about models, adapters, versions, storage paths, resource requirements, and statuses;
|
|
- does not store model binaries in git.
|
|
|
|
scripts/
|
|
- the operational center of the repo;
|
|
- contains validation, smoke tests, deployment scripts, reporting, indexing, model download helpers, 1C tooling, and service control utilities.
|
|
|
|
docs/
|
|
- runbooks, architecture notes, roadmap, API contracts, and research notes.
|
|
|
|
Current architectural principle:
|
|
- plugins may depend on core;
|
|
- core must not depend on plugins.
|
|
|
|
==================================================
|
|
3. PLUGINS OVERVIEW
|
|
==================================================
|
|
|
|
Text plugin:
|
|
- general text/chat/code-style use cases;
|
|
- currently supported by local model registry and vLLM deployment patterns.
|
|
|
|
Translation plugin:
|
|
- local translation service route exists;
|
|
- transformers-based service deployment is prepared.
|
|
|
|
Audio plugin:
|
|
- speech-related plugin;
|
|
- service route and deployment assets exist;
|
|
- smoke and readiness tooling exist.
|
|
|
|
Video plugin:
|
|
- vision/video analysis direction;
|
|
- service route and deployment assets exist;
|
|
- intended for frame/video understanding tasks.
|
|
|
|
Image plugin:
|
|
- generation and editing;
|
|
- SDXL is the practical current image route;
|
|
- Qwen image edit experimentation exists, but is much heavier/slower in practice.
|
|
|
|
1C plugin:
|
|
- the most mature and strategically important plugin in the repository;
|
|
- includes RAG, metadata parsing, BSL/module/form analysis, adapter contracts, connector policies, an agent service, training artifacts, evals, and many safety checks.
|
|
|
|
==================================================
|
|
4. THE 1C DIRECTION: WHY IT MATTERS
|
|
==================================================
|
|
|
|
The 1C plugin is the deepest part of the project. It is not a simple prompt layer. It is evolving into a safe assistant stack for 1C development and analysis.
|
|
|
|
What the 1C plugin is meant to do:
|
|
- answer 1C and BSL questions;
|
|
- help analyze metadata and object structure;
|
|
- help with read-only 1C queries;
|
|
- support RAG over 1C documentation and internal knowledge;
|
|
- inspect forms, modules, templates, and related artifacts;
|
|
- support controlled change planning;
|
|
- eventually support fine-tuned adapters/LoRA after enough high-quality examples are collected.
|
|
|
|
Important strategic rule:
|
|
- first RAG and tooling;
|
|
- fine-tuning later.
|
|
|
|
This is a deliberate choice. The project is trying to avoid premature fine-tuning before having enough verified, safe, domain-correct examples.
|
|
|
|
==================================================
|
|
5. WHAT HAS ALREADY BEEN IMPLEMENTED
|
|
==================================================
|
|
|
|
This repository is already beyond the "empty skeleton" stage. It has real operational substance.
|
|
|
|
Implemented foundation:
|
|
- architecture and repository layout;
|
|
- model registry with model cards;
|
|
- plugin structure for all target domains;
|
|
- deployment assets for GPU-hosted inference services;
|
|
- local model chat UI and service control patterns;
|
|
- large collection of validation and smoke scripts;
|
|
- runbooks for deployment and operation.
|
|
|
|
Implemented model/platform side:
|
|
- model registry in registry/model-cards;
|
|
- plugin model bundle in plugins/model-bundle.yaml;
|
|
- vLLM deployment assets;
|
|
- llama.cpp deployment assets;
|
|
- transformers-based deployment assets for translation/audio/video/image;
|
|
- runtime profile and GPU profile configuration;
|
|
- scripts for model download, validation, indexing, status collection, and service management.
|
|
|
|
Implemented UX/operations side:
|
|
- model chat server and UI for manual model checks;
|
|
- management console web assets;
|
|
- platform status collection/reporting;
|
|
- image generation/edit job flows and report artifacts;
|
|
- service control endpoints for starting/stopping model-serving stacks.
|
|
|
|
Implemented 1C side:
|
|
- 1C metadata parsing modules;
|
|
- XML/form/payload/dbnames/config-related parsing utilities;
|
|
- 1C RAG corpus preparation and lexical/vector index scripts;
|
|
- 1C prompts and RAG profile routing;
|
|
- 1C connector and adapter service boundaries;
|
|
- 1C MCP adapter code;
|
|
- 1C agent server with its own web UI;
|
|
- policy files for read-only queries and change workflow;
|
|
- schema files for metadata snapshots, BSL snapshots, and moxel registry;
|
|
- training example scaffolding and LoRA config scaffolding;
|
|
- many focused checks, smoke tests, analysis scripts, and verification scripts.
|
|
|
|
Evidence of maturity:
|
|
- about 267 files in scripts/;
|
|
- about 20 test files under tests/1c alone;
|
|
- extensive runbook documentation;
|
|
- many contract-first docs for 1C behavior and write safety.
|
|
|
|
==================================================
|
|
6. CURRENT STATUS AS OF THE DOCUMENTED ROADMAP
|
|
==================================================
|
|
|
|
The roadmap file contains a concrete status snapshot dated 2026-06-20. Treat this as the latest documented status inside the repo unless newer evidence is added elsewhere.
|
|
|
|
Documented operating status at that point:
|
|
- main UI available on LAN;
|
|
- current GPU profile was "image";
|
|
- model chat UI container was running;
|
|
- image and translation services were online;
|
|
- text, audio, video, and some 1C-heavy routes were prepared but not always running at the same time;
|
|
- the platform intentionally used a single-heavy-model approach on RTX 4090 to avoid VRAM conflicts.
|
|
|
|
Documented plugin state:
|
|
- Text: ready to start via vLLM route.
|
|
- 1C: ready to start, with a strong GGUF route preferred for manual checks.
|
|
- Translation: online.
|
|
- Audio: ready to start.
|
|
- Video: ready to start.
|
|
- Image: online, SDXL verified.
|
|
|
|
Documented completed priorities in roadmap:
|
|
- first-class launcher for Qwen3-Coder Q6 on GPU;
|
|
- 1C route preference for the strongest practical GGUF model for manual checks;
|
|
- smoke test additions for more profiles;
|
|
- image model mode switching;
|
|
- status report generation from health endpoints.
|
|
|
|
Practical interpretation:
|
|
- the platform already has operational deployment logic;
|
|
- however, not every route is meant to be hot simultaneously;
|
|
- the operator workflow and profile switching are part of the design, not a temporary bug.
|
|
|
|
==================================================
|
|
7. MODEL STRATEGY
|
|
==================================================
|
|
|
|
The repository uses a curated "one practical model per plugin" strategy rather than trying to host every possible large model.
|
|
|
|
Examples from the bundle:
|
|
- text: Qwen3-4B-Instruct;
|
|
- translation: LMT-60-4B;
|
|
- audio: Whisper Large V3 Turbo;
|
|
- video: Qwen2.5-VL 7B;
|
|
- image: SDXL Base plus SDXL Inpainting;
|
|
- 1C: Qwen3-Coder GGUF variant as a practical code/1C route.
|
|
|
|
Why this matters:
|
|
- the project is optimizing for practical local deployment;
|
|
- VRAM and runtime footprint are first-class constraints;
|
|
- model selection is tied to plugin purpose, deployment realism, and operator workflow.
|
|
|
|
==================================================
|
|
8. 1C SAFETY MODEL
|
|
==================================================
|
|
|
|
This is one of the most important parts for any AI reading this project.
|
|
|
|
The 1C direction is safety-first.
|
|
|
|
The model may:
|
|
- inspect metadata;
|
|
- search/read modules;
|
|
- validate read-only queries;
|
|
- propose changes;
|
|
- prepare plans and evidence;
|
|
- help build reviewable artifacts.
|
|
|
|
The model must not:
|
|
- directly change a live 1C database;
|
|
- run destructive queries;
|
|
- invent metadata when the adapter/connector does not confirm it;
|
|
- treat effective runtime view as a directly writable surface.
|
|
|
|
Core safety principles already documented in the repo:
|
|
- read-only first;
|
|
- full semantic 1C paths are preferred over storage internals;
|
|
- effective read view and editor provenance must stay separate;
|
|
- every write must begin with a plan;
|
|
- direct active configuration writes are forbidden;
|
|
- saved-state and extension-aware workflows are preferred;
|
|
- concrete write targets, guards, and validation are mandatory.
|
|
|
|
This means the 1C system is being designed more like a controlled engineering assistant than a free-form coding bot.
|
|
|
|
==================================================
|
|
9. WHAT IS IMPORTANT ABOUT THE 1C ADAPTER/CONNECTOR DESIGN
|
|
==================================================
|
|
|
|
The repo already contains a detailed contract for how 1C interaction should work.
|
|
|
|
High-level design ideas:
|
|
- the adapter works through SQL storage and controlled analysis layers;
|
|
- XML exports and Form.xml are useful for analysis and learning, but are not the live write transport;
|
|
- agent-facing APIs should speak in full 1C semantic paths, not raw SQL/CAS/internal offsets;
|
|
- metadata path resolution and BSL symbol resolution are separate problems;
|
|
- the adapter must preserve origin/provenance information so we always know whether something belongs to base config, extension, or saved-state.
|
|
|
|
Planned/implemented adapter-side capabilities include:
|
|
- object resolution;
|
|
- fact resolution;
|
|
- BSL symbol resolution;
|
|
- object context retrieval;
|
|
- modules/form/template inspection;
|
|
- write planning;
|
|
- saved-state targeting and verification;
|
|
- evidence-preserving decode behavior;
|
|
- smoke checks and contract checks.
|
|
|
|
This is already a serious design effort, not just a TODO note.
|
|
|
|
==================================================
|
|
10. WHAT REMAINS INCOMPLETE OR STILL EVOLVING
|
|
==================================================
|
|
|
|
Even though a lot exists, the project is still in an active build-out phase.
|
|
|
|
Areas that are clearly still evolving:
|
|
- final extraction of heavy plugins into standalone services;
|
|
- broader production-grade monitoring and lifecycle management;
|
|
- full maturity of video/audio/image routes under routine operation;
|
|
- richer model cabinet / ingest / verification workflows;
|
|
- more complete 1C live integration loops;
|
|
- broader eval coverage across plugins;
|
|
- final training/fine-tuning loop for 1C after verified datasets mature;
|
|
- more end-to-end workflows that turn read/analysis into safely validated change pipelines.
|
|
|
|
The 1C write path especially is intentionally conservative:
|
|
- planning and verification are much further along than unrestricted applying;
|
|
- the project prefers blocked/refused actions over unsafe hidden behavior;
|
|
- saved-state and extension workflows are being formalized step by step.
|
|
|
|
==================================================
|
|
11. WHAT WE HAVE DONE VS WHAT WE STILL NEED TO DO
|
|
==================================================
|
|
|
|
Already done:
|
|
- defined the platform direction and architecture;
|
|
- created the plugin structure;
|
|
- built the model registry concept and model-card inventory;
|
|
- prepared deployment stacks for local/GPU model serving;
|
|
- built a manual model-chat/testing UI;
|
|
- created many operational scripts and checks;
|
|
- established the 1C plugin as the main advanced domain;
|
|
- implemented large portions of 1C parsing, RAG, adapter contracts, and safety policies;
|
|
- documented many workflows and contracts in runbooks and design docs.
|
|
|
|
Still need to do:
|
|
- continue hardening the platform from "engineering bench" into "repeatable service platform";
|
|
- improve standalone service boundaries for heavy plugins;
|
|
- keep refining GPU profile switching and operational ergonomics;
|
|
- expand eval coverage and regression checks for all plugins;
|
|
- continue building the 1C live integration loop with strong provenance and write safety;
|
|
- improve extension-aware and saved-state-aware write planning and validation;
|
|
- grow verified 1C datasets;
|
|
- train and compare 1C-specific LoRA/adapters only after enough trustworthy examples exist;
|
|
- keep all dangerous operations gated by explicit planning, validation, and rollback-capable flows.
|
|
|
|
==================================================
|
|
12. RECOMMENDED MENTAL MODEL FOR ANOTHER AI
|
|
==================================================
|
|
|
|
If you are another AI reading this file, understand the project like this:
|
|
|
|
This repository is a local AI platform with a strong operations and tooling focus.
|
|
It is not only about prompting models. It is about:
|
|
- selecting the right local models;
|
|
- deploying them on constrained GPU hardware;
|
|
- routing use cases through plugin-specific services;
|
|
- verifying quality with smoke/eval checks;
|
|
- building safe interfaces to real-world systems.
|
|
|
|
The 1C area is the flagship domain.
|
|
Its current priority is safe read/analysis/RAG/tooling with controlled change planning.
|
|
Unsafe convenience is not acceptable there.
|
|
|
|
You should assume:
|
|
- reproducibility matters;
|
|
- deployment realism matters;
|
|
- model size/VRAM tradeoffs matter;
|
|
- contracts and runbooks matter;
|
|
- safety gates matter more than agent autonomy in 1C live workflows.
|
|
|
|
==================================================
|
|
13. HOW TO CONTRIBUTE CORRECTLY
|
|
==================================================
|
|
|
|
When continuing work in this repo, prefer these behaviors:
|
|
- preserve the core + plugins boundary;
|
|
- keep model binaries and large datasets out of git;
|
|
- add or update model cards instead of hardcoding assumptions;
|
|
- prefer operational scripts and reproducible runbooks over one-off manual steps;
|
|
- for 1C, prefer read-only analysis and explicit planning before any write path work;
|
|
- preserve provenance/origin information;
|
|
- use full semantic 1C paths where possible;
|
|
- do not collapse safe abstractions into raw storage details in agent-facing flows;
|
|
- add checks/tests/docs together with new behavior;
|
|
- treat the scripts/ folder as part of the product, not as disposable glue.
|
|
|
|
==================================================
|
|
14. MOST IMPORTANT NEXT STEPS
|
|
==================================================
|
|
|
|
The most sensible next steps, based on the repository state, are:
|
|
|
|
1. Continue strengthening the 1C safe operational loop.
|
|
- Better write planning.
|
|
- Better target resolution.
|
|
- Better saved-state verification.
|
|
- Better extension provenance and conflict detection.
|
|
|
|
2. Improve end-to-end readiness of non-1C plugins.
|
|
- Make text/audio/video/image routes easier to operate repeatedly.
|
|
- Expand smoke tests and health reporting.
|
|
- Reduce ambiguity in runtime profiles and active service state.
|
|
|
|
3. Keep the model registry and deployment inventory clean.
|
|
- Ensure model cards, bundles, and runtime profiles stay aligned.
|
|
- Keep practical model choices explicit.
|
|
|
|
4. Build confidence through validation.
|
|
- Prefer contract checks, smoke checks, and generated reports.
|
|
- Keep adding regression coverage where workflows are safety-sensitive.
|
|
|
|
5. Delay aggressive fine-tuning until the data is worthy.
|
|
- RAG and tooling first.
|
|
- Verified examples next.
|
|
- LoRA/adapters after that.
|
|
|
|
==================================================
|
|
15. SHORT VERSION
|
|
==================================================
|
|
|
|
We are building a self-hosted local LLM platform with a plugin architecture.
|
|
The platform supports text, translation, audio, video, image, and especially 1C.
|
|
|
|
What is already done:
|
|
- architecture;
|
|
- model registry;
|
|
- deployment stacks;
|
|
- chat/testing UI;
|
|
- large operations script layer;
|
|
- strong 1C parsing/RAG/adapter/safety foundation.
|
|
|
|
What is most important now:
|
|
- mature the platform operationally;
|
|
- keep non-1C plugins practical to run;
|
|
- continue deepening the 1C assistant with strict safety, provenance, planning, and validation;
|
|
- only move to 1C fine-tuning after enough verified domain data exists.
|
|
|
|
End of file.
|