# Model Chat UI Service Цель: поднять личный кабинет проверки моделей в локальной сети. LAN URL: ```text http://192.168.220.91:8765/tools/model-chat/ ``` ## Deploy Preflight before deployment: ```powershell powershell -ExecutionPolicy Bypass -File scripts\preflight_model_chat_ui.ps1 ` -DockerHost ssh://docker-gpu ` -SshTarget docker-gpu ` -CheckLive ``` Fast local-only preflight: ```powershell powershell -ExecutionPolicy Bypass -File scripts\preflight_model_chat_ui.ps1 -SkipDockerConfig ``` ```powershell powershell -ExecutionPolicy Bypass -File scripts/deploy_model_chat_ui.ps1 ` -DockerHost ssh://docker-gpu ` -SshTarget docker-gpu ``` Скрипт: - собирает небольшой архив приложения без тяжелых `models/`; - копирует его на Windows GPU-хост; - распаковывает в `Z:\LLM\model-chat-app`; - запускает compose-сервис `llm-model-chat-ui`; - открывает Windows Firewall для TCP `8765`. ## Health ```powershell Invoke-RestMethod http://192.168.220.91:8765/api/health Invoke-RestMethod http://192.168.220.91:8765/api/catalog Invoke-RestMethod http://192.168.220.91:8765/api/model-services python scripts\plan_model_services.py python scripts\generate_model_chat_status.py ``` The status generator writes a compact Markdown snapshot to: ```text reports/model-chat/status.md ``` `/api/health` includes: - endpoint status for vLLM, llama.cpp, translation, audio, and video service ports; - selected route for each plugin; - GPU profile readiness with missing services to start and conflicting services to stop; - latest Model Chat UI preflight status from `reports/model-chat/preflight.json`; - local storage status for every registered model; - service plan with compose/deploy hints. Service states: - `online`: endpoint is running and reports the expected served model name; - `ready_to_start`: model files are present, but the service is not online; - `blocked`: local model files are missing or incomplete. ## Service Control The UI has a `service` panel for the selected model. It calls `POST /api/service-control`: ```json { "model_id": "whisper-large-v3-turbo", "action": "status" } ``` Allowed actions are `start`, `stop`, `restart`, and `status`. If the UI container has no Docker CLI or SSH key, use the operator script from the project folder: ```powershell powershell -ExecutionPolicy Bypass -File scripts\manage_model_service.ps1 -Service audio-api -Action status powershell -ExecutionPolicy Bypass -File scripts\manage_model_service.ps1 -Service audio-api -Action stop powershell -ExecutionPolicy Bypass -File scripts\manage_model_service.ps1 -Service translation-api -Action start ``` Use `stop` on heavy services before starting another large model when VRAM is low. ## GPU Profiles Use GPU profiles instead of manual container juggling when switching between heavy models: ```powershell powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile default powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile text powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile audio powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile video powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile image powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile gguf-1c ``` Profile definitions live in: ```text config/gpu_profiles.json ``` The UI catalog, `/api/health` profile readiness, and `scripts\switch_gpu_profile.ps1` read the same file. Validate it before deploy: ```powershell python scripts\validate_gpu_profiles.py --print ``` Validate the UI deployment archive without uploading it: ```powershell powershell -ExecutionPolicy Bypass -File scripts\deploy_model_chat_ui.ps1 -ArchiveOnly ``` Dry-run the plan without touching containers: ```powershell powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile video -PlanOnly ``` Profiles: - `default` / `text`: vLLM text, translation, audio, and UI; video, image, and llama.cpp stopped. - `audio`: audio, translation, vLLM text, and UI; video, image, and llama.cpp stopped. - `video`: translation, video, and UI; vLLM text, audio, image, and llama.cpp stopped to free VRAM. - `image`: translation, image, and UI; vLLM text, audio, video, and llama.cpp stopped to free VRAM. - `gguf-1c`: llama.cpp, translation, and UI; vLLM text, audio, video, and image stopped. The Model Chat UI also shows the recommended profile command for the selected plugin. The profile hint shows `status: ready` when the current containers already match the profile. Otherwise it lists services that should be started and services that should be stopped. For `video` and `image`, the service starts quickly but the first request can still spend several minutes loading the model into GPU memory. ## Runtime Profiles Runtime profiles choose the execution host and endpoint used by the chat UI: ```text config/runtime_profiles.json ``` Profiles: - `gpu-fast`: default interactive profile on `docker-gpu.cin.su`; uses RTX 4090 endpoints. - `cpu-test`: benchmark/fallback profile on `docker-test.cin.su`; currently maps Qwen3-Coder Q6 to `http://docker-test.cin.su:18086` with served model `qwen3-coder-1c-q6-cpu`. - `background`: batch profile for downloads, RAG indexing and conversions; not intended for direct chat. The UI applies `model_overrides` from the selected runtime profile. This allows the same registry model to use a different endpoint or served model name on another host. The service panel has `Benchmark GPU / CPU`. It calls `POST /api/benchmark/runtime`, runs `scripts/benchmark_runtime_profiles.py`, and stores the raw report in: ```text /reports/benchmarks/runtime-profiles--.json ``` The UI loads recent benchmark history for the currently selected model/plugin through `GET /api/benchmark/history?model_id=&plugin=`. When both `gpu-fast` and `cpu-test` succeed, the report includes `speedup.gpu_vs_cpu_ratio`. Use larger generation limits, for example 192-384 tokens, for representative GPU/CPU ratios; very short runs include more startup and request overhead. The UI has a separate `Benchmark tokens` field, default `384`, so chat generation limits do not accidentally make CPU comparison runs too long. Before starting the long benchmark request, the UI performs a quick preflight for both `gpu-fast` and `cpu-test`; if either served model is missing, the benchmark is not started. ```text POST /api/benchmark/preflight ``` CLI check: ```powershell powershell -NoProfile -ExecutionPolicy Bypass -File scripts\check_runtime_preflight.ps1 ``` Each history row links to the raw JSON report through: ```text /api/benchmark/report?name= ``` It also links to a generated Markdown summary: ```text /api/benchmark/report.md?name= ``` For `cpu-test`, start the heavyweight llama.cpp CPU server before sending chat requests: ```powershell powershell -ExecutionPolicy Bypass -File scripts\manage_cpu_q6_service.ps1 -Action start powershell -ExecutionPolicy Bypass -File scripts\manage_cpu_q6_service.ps1 -Action status powershell -ExecutionPolicy Bypass -File scripts\manage_cpu_q6_service.ps1 -Action logs ``` For `gpu-fast` Qwen3-Coder Q6 checks on port `8081`, use the GPU launcher. Stop image/video/vLLM first if VRAM is tight: ```powershell powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile gguf-1c powershell -ExecutionPolicy Bypass -File scripts\manage_gpu_q6_service.ps1 -Action start powershell -ExecutionPolicy Bypass -File scripts\manage_gpu_q6_service.ps1 -Action status powershell -ExecutionPolicy Bypass -File scripts\manage_gpu_q6_service.ps1 -Action logs ``` After you convert a trained PEFT LoRA adapter to GGUF for `llama.cpp`, you can start the same service with the adapter applied: ```powershell powershell -ExecutionPolicy Bypass -File scripts\convert_1c_lora_to_gguf_gpu.ps1 powershell -ExecutionPolicy Bypass -File scripts\manage_gpu_q6_service.ps1 ` -Action start ` -LoraPath /models/adapters/1c/qwen3-coder-30b-a3b-1c-lora-v1.gguf ``` Stop it after benchmarks on the shared host if it is no longer needed: ```powershell powershell -ExecutionPolicy Bypass -File scripts\manage_cpu_q6_service.ps1 -Action stop powershell -ExecutionPolicy Bypass -File scripts\manage_gpu_q6_service.ps1 -Action stop ``` ## Containers ```powershell docker -H ssh://docker-gpu ps docker -H ssh://docker-gpu logs --tail 100 llm-model-chat-ui docker -H ssh://docker-gpu logs --tail 100 llm-vllm-text docker -H ssh://docker-gpu logs --tail 100 llm-transformers-translation ``` ## Transformers Plugin Services Translation, audio, video, and image services use the small OpenAI-like server in `scripts/transformers_plugin_server.py`. ```powershell powershell -ExecutionPolicy Bypass -File scripts\deploy_transformers_service.ps1 -Plugin translation powershell -ExecutionPolicy Bypass -File scripts\deploy_transformers_service.ps1 -Plugin audio powershell -ExecutionPolicy Bypass -File scripts\deploy_transformers_service.ps1 -Plugin video powershell -ExecutionPolicy Bypass -File scripts\deploy_transformers_service.ps1 -Plugin image ``` By default the deploy script also refreshes `Z:\LLM\model-chat-app` before restarting the selected service, so changes in `scripts/transformers_plugin_server.py` are applied to the mounted `/app` directory. Use `-NoSyncApp` only for a fast container restart when the app files are already current. Ports: - translation: `8010` - audio: `8020` - video: `8030` - image: `8040` The translation service exposes `/health`, `/v1/models`, and `/v1/chat/completions`. The audio service exposes `/health`, `/v1/models`, and `/v1/audio/transcriptions`; in the UI, select plugin `Звук`, choose an audio file, then press `Отправить`. The video service exposes `/health`, `/v1/models`, and `/v1/vision/analyze`; in the UI, select plugin `Видео`, choose an image, enter the question in the prompt box, then press `Отправить`. The image service exposes `/health`, `/v1/models`, `/v1/images/generations`, `/v1/images/edits`, and async job endpoints under `/v1/images/jobs`. In the UI, select plugin `Фото`, choose an image model, choose `generate` or `edit / inpaint`, enter the prompt, then press `Отправить`. By default `LOAD_ON_START=1` and `BACKGROUND_LOAD_ON_START=1` for image generation, so the service opens HTTP quickly and loads the SDXL base model in the background. The first request can show `loading_model` for several minutes. Switching between SDXL base and inpaint may unload the other pipeline to keep VRAM available. Stop the large text vLLM container before loading another large model if VRAM is tight. The UI can select `Qwen Image Edit`, but the image endpoint must actually serve `qwen-image-edit`; otherwise `/api/image/submit` rejects the job instead of silently using SDXL. To test Qwen Image Edit as a single-heavy-model experiment, restart the image service with the Qwen env file: ```powershell docker -H ssh://docker-gpu compose ` --env-file core/deploy/docker-gpu/transformers/image.qwen-edit.env.example ` -f core/deploy/docker-gpu/transformers/image.compose.yaml up -d ``` Restore the verified SDXL service with: ```powershell powershell -ExecutionPolicy Bypass -File scripts\deploy_transformers_service.ps1 -Plugin image ``` Verified Qwen Image Edit notes from `2026-06-20`: - model files are present under `Z:\LLM\models\image\qwen-image-edit`; - `diffusers` detects `QwenImageEditPipeline` and the service can expose `/v1/models` as `qwen-image-edit`; - pipeline load completed in about `31 s` with CPU offload on RTX 4090; - a 512x512 edit job with `1` inference step did not finish within `1800 s`, so Qwen Image Edit is not practical on the current 24 GB GPU profile without a quantized/optimized runtime or a larger GPU; - keep SDXL as the default verified image service for now. Smoke test image jobs through the Model Chat proxy: ```powershell python scripts\smoke_image_jobs.py --operation generate --steps 4 python scripts\smoke_image_jobs.py --operation edit --steps 4 python scripts\smoke_image_jobs.py --operation edit --steps 40 --cancel --cancel-after 1 python scripts\smoke_image_jobs.py --model-id qwen-image-edit --model-mode qwen-image-edit --model qwen-image-edit --operation edit --steps 4 ``` Audio and video smoke checks generate small synthetic inputs locally and send them through the Model Chat UI proxy: ```powershell powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile audio python scripts\smoke_audio_transcription.py powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile video python scripts\smoke_video_analysis.py ``` When the matching service is intentionally stopped, use `--allow-unavailable` for a code-path check that does not fail the local validation run. Verified on `2026-06-20` with RTX 4090: - generate, SDXL base, 512x512, 4 steps: completed in `2194 ms` after warmup; - edit / inpaint, SDXL inpaint, 512x512, 4 steps: completed in `180649 ms` including first inpaint model load; - generate, SDXL base, 512x512, 1 step: completed in `86866 ms` after image service restart and warmup; - generated artifacts are served through `/generated-images//.png`. Direct artifact check uses `GET`; this minimal HTTP server does not implement `HEAD` for generated files. ## Model Downloads On GPU Host Use this when model files should be written directly to `Z:\LLM\models` on the GPU host: For large Hugging Face shard files, prefer the explicit range downloader. It writes final files directly to `/models/...`, resumes by local file size, and avoids stale `.cache/huggingface/download/*.incomplete` files left by interrupted Xet/snapshot downloads. ```powershell powershell -ExecutionPolicy Bypass -File scripts\download_hf_range_gpu.ps1 ` -CardId qwen2_5-vl-7b-instruct ` -LocalDir /models/video/qwen2.5-vl-7b-instruct ` -AllowFile "model-00001-of-00005.safetensors,model-00002-of-00005.safetensors,model-00003-of-00005.safetensors,model-00004-of-00005.safetensors,model-00005-of-00005.safetensors" ` -Detached ` -ContainerName llm-hf-range-qwen-vl ` -ChunkSize 256mb ``` The registry/snapshot downloader is still useful for dry-run planning and small metadata files: ```powershell powershell -ExecutionPolicy Bypass -File scripts\download_missing_hf_models_gpu.ps1 -DryRun docker -H ssh://docker-gpu rm -f llm-model-download powershell -ExecutionPolicy Bypass -File scripts\download_missing_hf_models_gpu.ps1 ` -Detached ` -ContainerName llm-model-download ``` Monitor: ```powershell docker -H ssh://docker-gpu ps docker -H ssh://docker-gpu logs --tail 100 llm-model-download Invoke-RestMethod http://192.168.220.91:8765/api/model-services ``` ## Notes - UI is served by `scripts/model_chat_server.py`. - Inference endpoint defaults to `http://docker-gpu.cin.su:8000`. - Use `vllm/vllm-openai:v0.10.2` for the UI image because it already contains Python and is verified on this host. - `vllm/vllm-openai:latest` passes CUDA after driver `595.97`, but vLLM `0.23.0` currently fails on Docker Desktop/WSL with `UVA is not available`.