Files
llm/docs/runbooks/model-chat-ui-service.md

15 KiB

Model Chat UI Service

Цель: поднять личный кабинет проверки моделей в локальной сети.

LAN URL:

http://192.168.220.91:8765/tools/model-chat/

Deploy

Preflight before deployment:

powershell -ExecutionPolicy Bypass -File scripts\preflight_model_chat_ui.ps1 `
  -DockerHost ssh://docker-gpu `
  -SshTarget docker-gpu `
  -CheckLive

Fast local-only preflight:

powershell -ExecutionPolicy Bypass -File scripts\preflight_model_chat_ui.ps1 -SkipDockerConfig
powershell -ExecutionPolicy Bypass -File scripts/deploy_model_chat_ui.ps1 `
  -DockerHost ssh://docker-gpu `
  -SshTarget docker-gpu

Скрипт:

  • собирает небольшой архив приложения без тяжелых models/;
  • копирует его на Windows GPU-хост;
  • распаковывает в Z:\LLM\model-chat-app;
  • запускает compose-сервис llm-model-chat-ui;
  • открывает Windows Firewall для TCP 8765.

Health

Invoke-RestMethod http://192.168.220.91:8765/api/health
Invoke-RestMethod http://192.168.220.91:8765/api/catalog
Invoke-RestMethod http://192.168.220.91:8765/api/model-services
python scripts\plan_model_services.py
python scripts\generate_model_chat_status.py

The status generator writes a compact Markdown snapshot to:

reports/model-chat/status.md

/api/health includes:

  • endpoint status for vLLM, llama.cpp, translation, audio, and video service ports;
  • selected route for each plugin;
  • GPU profile readiness with missing services to start and conflicting services to stop;
  • latest Model Chat UI preflight status from reports/model-chat/preflight.json;
  • local storage status for every registered model;
  • service plan with compose/deploy hints.

Service states:

  • online: endpoint is running and reports the expected served model name;
  • ready_to_start: model files are present, but the service is not online;
  • blocked: local model files are missing or incomplete.

Service Control

The UI has a service panel for the selected model. It calls POST /api/service-control:

{
  "model_id": "whisper-large-v3-turbo",
  "action": "status"
}

Allowed actions are start, stop, restart, and status.

If the UI container has no Docker CLI or SSH key, use the operator script from the project folder:

powershell -ExecutionPolicy Bypass -File scripts\manage_model_service.ps1 -Service audio-api -Action status
powershell -ExecutionPolicy Bypass -File scripts\manage_model_service.ps1 -Service audio-api -Action stop
powershell -ExecutionPolicy Bypass -File scripts\manage_model_service.ps1 -Service translation-api -Action start

Use stop on heavy services before starting another large model when VRAM is low.

GPU Profiles

Use GPU profiles instead of manual container juggling when switching between heavy models:

powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile default
powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile text
powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile audio
powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile video
powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile image
powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile gguf-1c

Profile definitions live in:

config/gpu_profiles.json

The UI catalog, /api/health profile readiness, and scripts\switch_gpu_profile.ps1 read the same file. Validate it before deploy:

python scripts\validate_gpu_profiles.py --print

Validate the UI deployment archive without uploading it:

powershell -ExecutionPolicy Bypass -File scripts\deploy_model_chat_ui.ps1 -ArchiveOnly

Dry-run the plan without touching containers:

powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile video -PlanOnly

Profiles:

  • default / text: vLLM text, translation, audio, and UI; video, image, and llama.cpp stopped.
  • audio: audio, translation, vLLM text, and UI; video, image, and llama.cpp stopped.
  • video: translation, video, and UI; vLLM text, audio, image, and llama.cpp stopped to free VRAM.
  • image: translation, image, and UI; vLLM text, audio, video, and llama.cpp stopped to free VRAM.
  • gguf-1c: llama.cpp, translation, and UI; vLLM text, audio, video, and image stopped.

The Model Chat UI also shows the recommended profile command for the selected plugin. The profile hint shows status: ready when the current containers already match the profile. Otherwise it lists services that should be started and services that should be stopped. For video and image, the service starts quickly but the first request can still spend several minutes loading the model into GPU memory.

Runtime Profiles

Runtime profiles choose the execution host and endpoint used by the chat UI:

config/runtime_profiles.json

Profiles:

  • gpu-fast: default interactive profile on docker-gpu.cin.su; uses RTX 4090 endpoints.
  • cpu-test: benchmark/fallback profile on docker-test.cin.su; currently maps Qwen3-Coder Q6 to http://docker-test.cin.su:18086 with served model qwen3-coder-1c-q6-cpu.
  • background: batch profile for downloads, RAG indexing and conversions; not intended for direct chat.

The UI applies model_overrides from the selected runtime profile. This allows the same registry model to use a different endpoint or served model name on another host.

The service panel has Benchmark GPU / CPU. It calls POST /api/benchmark/runtime, runs scripts/benchmark_runtime_profiles.py, and stores the raw report in:

/reports/benchmarks/runtime-profiles-<model-id>-<timestamp>.json

The UI loads recent benchmark history for the currently selected model/plugin through GET /api/benchmark/history?model_id=<model-id>&plugin=<plugin-id>. When both gpu-fast and cpu-test succeed, the report includes speedup.gpu_vs_cpu_ratio. Use larger generation limits, for example 192-384 tokens, for representative GPU/CPU ratios; very short runs include more startup and request overhead. The UI has a separate Benchmark tokens field, default 384, so chat generation limits do not accidentally make CPU comparison runs too long. Before starting the long benchmark request, the UI performs a quick preflight for both gpu-fast and cpu-test; if either served model is missing, the benchmark is not started.

POST /api/benchmark/preflight

CLI check:

powershell -NoProfile -ExecutionPolicy Bypass -File scripts\check_runtime_preflight.ps1

Each history row links to the raw JSON report through:

/api/benchmark/report?name=<runtime-profiles-report.json>

It also links to a generated Markdown summary:

/api/benchmark/report.md?name=<runtime-profiles-report.json>

For cpu-test, start the heavyweight llama.cpp CPU server before sending chat requests:

powershell -ExecutionPolicy Bypass -File scripts\manage_cpu_q6_service.ps1 -Action start
powershell -ExecutionPolicy Bypass -File scripts\manage_cpu_q6_service.ps1 -Action status
powershell -ExecutionPolicy Bypass -File scripts\manage_cpu_q6_service.ps1 -Action logs

For gpu-fast Qwen3-Coder Q6 checks on port 8081, use the GPU launcher. Stop image/video/vLLM first if VRAM is tight:

powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile gguf-1c
powershell -ExecutionPolicy Bypass -File scripts\manage_gpu_q6_service.ps1 -Action start
powershell -ExecutionPolicy Bypass -File scripts\manage_gpu_q6_service.ps1 -Action status
powershell -ExecutionPolicy Bypass -File scripts\manage_gpu_q6_service.ps1 -Action logs

After you convert a trained PEFT LoRA adapter to GGUF for llama.cpp, you can start the same service with the adapter applied:

powershell -ExecutionPolicy Bypass -File scripts\convert_1c_lora_to_gguf_gpu.ps1

powershell -ExecutionPolicy Bypass -File scripts\manage_gpu_q6_service.ps1 `
  -Action start `
  -LoraPath /models/adapters/1c/qwen3-coder-30b-a3b-1c-lora-v1.gguf

Stop it after benchmarks on the shared host if it is no longer needed:

powershell -ExecutionPolicy Bypass -File scripts\manage_cpu_q6_service.ps1 -Action stop
powershell -ExecutionPolicy Bypass -File scripts\manage_gpu_q6_service.ps1 -Action stop

Containers

docker -H ssh://docker-gpu ps
docker -H ssh://docker-gpu logs --tail 100 llm-model-chat-ui
docker -H ssh://docker-gpu logs --tail 100 llm-vllm-text
docker -H ssh://docker-gpu logs --tail 100 llm-transformers-translation

Transformers Plugin Services

Translation, audio, video, and image services use the small OpenAI-like server in scripts/transformers_plugin_server.py.

powershell -ExecutionPolicy Bypass -File scripts\deploy_transformers_service.ps1 -Plugin translation
powershell -ExecutionPolicy Bypass -File scripts\deploy_transformers_service.ps1 -Plugin audio
powershell -ExecutionPolicy Bypass -File scripts\deploy_transformers_service.ps1 -Plugin video
powershell -ExecutionPolicy Bypass -File scripts\deploy_transformers_service.ps1 -Plugin image

By default the deploy script also refreshes Z:\LLM\model-chat-app before restarting the selected service, so changes in scripts/transformers_plugin_server.py are applied to the mounted /app directory. Use -NoSyncApp only for a fast container restart when the app files are already current.

Ports:

  • translation: 8010
  • audio: 8020
  • video: 8030
  • image: 8040

The translation service exposes /health, /v1/models, and /v1/chat/completions. The audio service exposes /health, /v1/models, and /v1/audio/transcriptions; in the UI, select plugin Звук, choose an audio file, then press Отправить. The video service exposes /health, /v1/models, and /v1/vision/analyze; in the UI, select plugin Видео, choose an image, enter the question in the prompt box, then press Отправить. The image service exposes /health, /v1/models, /v1/images/generations, /v1/images/edits, and async job endpoints under /v1/images/jobs. In the UI, select plugin Фото, choose an image model, choose generate or edit / inpaint, enter the prompt, then press Отправить. By default LOAD_ON_START=1 and BACKGROUND_LOAD_ON_START=1 for image generation, so the service opens HTTP quickly and loads the SDXL base model in the background. The first request can show loading_model for several minutes. Switching between SDXL base and inpaint may unload the other pipeline to keep VRAM available. Stop the large text vLLM container before loading another large model if VRAM is tight.

The UI can select Qwen Image Edit, but the image endpoint must actually serve qwen-image-edit; otherwise /api/image/submit rejects the job instead of silently using SDXL. To test Qwen Image Edit as a single-heavy-model experiment, restart the image service with the Qwen env file:

docker -H ssh://docker-gpu compose `
  --env-file core/deploy/docker-gpu/transformers/image.qwen-edit.env.example `
  -f core/deploy/docker-gpu/transformers/image.compose.yaml up -d

Restore the verified SDXL service with:

powershell -ExecutionPolicy Bypass -File scripts\deploy_transformers_service.ps1 -Plugin image

Verified Qwen Image Edit notes from 2026-06-20:

  • model files are present under Z:\LLM\models\image\qwen-image-edit;
  • diffusers detects QwenImageEditPipeline and the service can expose /v1/models as qwen-image-edit;
  • pipeline load completed in about 31 s with CPU offload on RTX 4090;
  • a 512x512 edit job with 1 inference step did not finish within 1800 s, so Qwen Image Edit is not practical on the current 24 GB GPU profile without a quantized/optimized runtime or a larger GPU;
  • keep SDXL as the default verified image service for now.

Smoke test image jobs through the Model Chat proxy:

python scripts\smoke_image_jobs.py --operation generate --steps 4
python scripts\smoke_image_jobs.py --operation edit --steps 4
python scripts\smoke_image_jobs.py --operation edit --steps 40 --cancel --cancel-after 1
python scripts\smoke_image_jobs.py --model-id qwen-image-edit --model-mode qwen-image-edit --model qwen-image-edit --operation edit --steps 4

Audio and video smoke checks generate small synthetic inputs locally and send them through the Model Chat UI proxy:

powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile audio
python scripts\smoke_audio_transcription.py

powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile video
python scripts\smoke_video_analysis.py

When the matching service is intentionally stopped, use --allow-unavailable for a code-path check that does not fail the local validation run.

Verified on 2026-06-20 with RTX 4090:

  • generate, SDXL base, 512x512, 4 steps: completed in 2194 ms after warmup;
  • edit / inpaint, SDXL inpaint, 512x512, 4 steps: completed in 180649 ms including first inpaint model load;
  • generate, SDXL base, 512x512, 1 step: completed in 86866 ms after image service restart and warmup;
  • generated artifacts are served through /generated-images/<date>/<file>.png.

Direct artifact check uses GET; this minimal HTTP server does not implement HEAD for generated files.

Model Downloads On GPU Host

Use this when model files should be written directly to Z:\LLM\models on the GPU host:

For large Hugging Face shard files, prefer the explicit range downloader. It writes final files directly to /models/..., resumes by local file size, and avoids stale .cache/huggingface/download/*.incomplete files left by interrupted Xet/snapshot downloads.

powershell -ExecutionPolicy Bypass -File scripts\download_hf_range_gpu.ps1 `
  -CardId qwen2_5-vl-7b-instruct `
  -LocalDir /models/video/qwen2.5-vl-7b-instruct `
  -AllowFile "model-00001-of-00005.safetensors,model-00002-of-00005.safetensors,model-00003-of-00005.safetensors,model-00004-of-00005.safetensors,model-00005-of-00005.safetensors" `
  -Detached `
  -ContainerName llm-hf-range-qwen-vl `
  -ChunkSize 256mb

The registry/snapshot downloader is still useful for dry-run planning and small metadata files:

powershell -ExecutionPolicy Bypass -File scripts\download_missing_hf_models_gpu.ps1 -DryRun

docker -H ssh://docker-gpu rm -f llm-model-download
powershell -ExecutionPolicy Bypass -File scripts\download_missing_hf_models_gpu.ps1 `
  -Detached `
  -ContainerName llm-model-download

Monitor:

docker -H ssh://docker-gpu ps
docker -H ssh://docker-gpu logs --tail 100 llm-model-download
Invoke-RestMethod http://192.168.220.91:8765/api/model-services

Notes

  • UI is served by scripts/model_chat_server.py.
  • Inference endpoint defaults to http://docker-gpu.cin.su:8000.
  • Use vllm/vllm-openai:v0.10.2 for the UI image because it already contains Python and is verified on this host.
  • vllm/vllm-openai:latest passes CUDA after driver 595.97, but vLLM 0.23.0 currently fails on Docker Desktop/WSL with UVA is not available.