15 KiB
Model Chat UI Service
Цель: поднять личный кабинет проверки моделей в локальной сети.
LAN URL:
http://192.168.220.91:8765/tools/model-chat/
Deploy
Preflight before deployment:
powershell -ExecutionPolicy Bypass -File scripts\preflight_model_chat_ui.ps1 `
-DockerHost ssh://docker-gpu `
-SshTarget docker-gpu `
-CheckLive
Fast local-only preflight:
powershell -ExecutionPolicy Bypass -File scripts\preflight_model_chat_ui.ps1 -SkipDockerConfig
powershell -ExecutionPolicy Bypass -File scripts/deploy_model_chat_ui.ps1 `
-DockerHost ssh://docker-gpu `
-SshTarget docker-gpu
Скрипт:
- собирает небольшой архив приложения без тяжелых
models/; - копирует его на Windows GPU-хост;
- распаковывает в
Z:\LLM\model-chat-app; - запускает compose-сервис
llm-model-chat-ui; - открывает Windows Firewall для TCP
8765.
Health
Invoke-RestMethod http://192.168.220.91:8765/api/health
Invoke-RestMethod http://192.168.220.91:8765/api/catalog
Invoke-RestMethod http://192.168.220.91:8765/api/model-services
python scripts\plan_model_services.py
python scripts\generate_model_chat_status.py
The status generator writes a compact Markdown snapshot to:
reports/model-chat/status.md
/api/health includes:
- endpoint status for vLLM, llama.cpp, translation, audio, and video service ports;
- selected route for each plugin;
- GPU profile readiness with missing services to start and conflicting services to stop;
- latest Model Chat UI preflight status from
reports/model-chat/preflight.json; - local storage status for every registered model;
- service plan with compose/deploy hints.
Service states:
online: endpoint is running and reports the expected served model name;ready_to_start: model files are present, but the service is not online;blocked: local model files are missing or incomplete.
Service Control
The UI has a service panel for the selected model. It calls POST /api/service-control:
{
"model_id": "whisper-large-v3-turbo",
"action": "status"
}
Allowed actions are start, stop, restart, and status.
If the UI container has no Docker CLI or SSH key, use the operator script from the project folder:
powershell -ExecutionPolicy Bypass -File scripts\manage_model_service.ps1 -Service audio-api -Action status
powershell -ExecutionPolicy Bypass -File scripts\manage_model_service.ps1 -Service audio-api -Action stop
powershell -ExecutionPolicy Bypass -File scripts\manage_model_service.ps1 -Service translation-api -Action start
Use stop on heavy services before starting another large model when VRAM is low.
GPU Profiles
Use GPU profiles instead of manual container juggling when switching between heavy models:
powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile default
powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile text
powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile audio
powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile video
powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile image
powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile gguf-1c
Profile definitions live in:
config/gpu_profiles.json
The UI catalog, /api/health profile readiness, and scripts\switch_gpu_profile.ps1 read the same file.
Validate it before deploy:
python scripts\validate_gpu_profiles.py --print
Validate the UI deployment archive without uploading it:
powershell -ExecutionPolicy Bypass -File scripts\deploy_model_chat_ui.ps1 -ArchiveOnly
Dry-run the plan without touching containers:
powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile video -PlanOnly
Profiles:
default/text: vLLM text, translation, audio, and UI; video, image, and llama.cpp stopped.audio: audio, translation, vLLM text, and UI; video, image, and llama.cpp stopped.video: translation, video, and UI; vLLM text, audio, image, and llama.cpp stopped to free VRAM.image: translation, image, and UI; vLLM text, audio, video, and llama.cpp stopped to free VRAM.gguf-1c: llama.cpp, translation, and UI; vLLM text, audio, video, and image stopped.
The Model Chat UI also shows the recommended profile command for the selected plugin.
The profile hint shows status: ready when the current containers already match the profile.
Otherwise it lists services that should be started and services that should be stopped.
For video and image, the service starts quickly but the first request can still spend several minutes loading
the model into GPU memory.
Runtime Profiles
Runtime profiles choose the execution host and endpoint used by the chat UI:
config/runtime_profiles.json
Profiles:
gpu-fast: default interactive profile ondocker-gpu.cin.su; uses RTX 4090 endpoints.cpu-test: benchmark/fallback profile ondocker-test.cin.su; currently maps Qwen3-Coder Q6 tohttp://docker-test.cin.su:18086with served modelqwen3-coder-1c-q6-cpu.background: batch profile for downloads, RAG indexing and conversions; not intended for direct chat.
The UI applies model_overrides from the selected runtime profile. This allows the same registry model
to use a different endpoint or served model name on another host.
The service panel has Benchmark GPU / CPU. It calls POST /api/benchmark/runtime, runs
scripts/benchmark_runtime_profiles.py, and stores the raw report in:
/reports/benchmarks/runtime-profiles-<model-id>-<timestamp>.json
The UI loads recent benchmark history for the currently selected model/plugin through
GET /api/benchmark/history?model_id=<model-id>&plugin=<plugin-id>.
When both gpu-fast and cpu-test succeed, the report includes speedup.gpu_vs_cpu_ratio.
Use larger generation limits, for example 192-384 tokens, for representative GPU/CPU ratios;
very short runs include more startup and request overhead.
The UI has a separate Benchmark tokens field, default 384, so chat generation limits do not
accidentally make CPU comparison runs too long.
Before starting the long benchmark request, the UI performs a quick preflight for both gpu-fast
and cpu-test; if either served model is missing, the benchmark is not started.
POST /api/benchmark/preflight
CLI check:
powershell -NoProfile -ExecutionPolicy Bypass -File scripts\check_runtime_preflight.ps1
Each history row links to the raw JSON report through:
/api/benchmark/report?name=<runtime-profiles-report.json>
It also links to a generated Markdown summary:
/api/benchmark/report.md?name=<runtime-profiles-report.json>
For cpu-test, start the heavyweight llama.cpp CPU server before sending chat requests:
powershell -ExecutionPolicy Bypass -File scripts\manage_cpu_q6_service.ps1 -Action start
powershell -ExecutionPolicy Bypass -File scripts\manage_cpu_q6_service.ps1 -Action status
powershell -ExecutionPolicy Bypass -File scripts\manage_cpu_q6_service.ps1 -Action logs
For gpu-fast Qwen3-Coder Q6 checks on port 8081, use the GPU launcher. Stop image/video/vLLM
first if VRAM is tight:
powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile gguf-1c
powershell -ExecutionPolicy Bypass -File scripts\manage_gpu_q6_service.ps1 -Action start
powershell -ExecutionPolicy Bypass -File scripts\manage_gpu_q6_service.ps1 -Action status
powershell -ExecutionPolicy Bypass -File scripts\manage_gpu_q6_service.ps1 -Action logs
After you convert a trained PEFT LoRA adapter to GGUF for llama.cpp, you can start the same
service with the adapter applied:
powershell -ExecutionPolicy Bypass -File scripts\convert_1c_lora_to_gguf_gpu.ps1
powershell -ExecutionPolicy Bypass -File scripts\manage_gpu_q6_service.ps1 `
-Action start `
-LoraPath /models/adapters/1c/qwen3-coder-30b-a3b-1c-lora-v1.gguf
Stop it after benchmarks on the shared host if it is no longer needed:
powershell -ExecutionPolicy Bypass -File scripts\manage_cpu_q6_service.ps1 -Action stop
powershell -ExecutionPolicy Bypass -File scripts\manage_gpu_q6_service.ps1 -Action stop
Containers
docker -H ssh://docker-gpu ps
docker -H ssh://docker-gpu logs --tail 100 llm-model-chat-ui
docker -H ssh://docker-gpu logs --tail 100 llm-vllm-text
docker -H ssh://docker-gpu logs --tail 100 llm-transformers-translation
Transformers Plugin Services
Translation, audio, video, and image services use the small OpenAI-like server in
scripts/transformers_plugin_server.py.
powershell -ExecutionPolicy Bypass -File scripts\deploy_transformers_service.ps1 -Plugin translation
powershell -ExecutionPolicy Bypass -File scripts\deploy_transformers_service.ps1 -Plugin audio
powershell -ExecutionPolicy Bypass -File scripts\deploy_transformers_service.ps1 -Plugin video
powershell -ExecutionPolicy Bypass -File scripts\deploy_transformers_service.ps1 -Plugin image
By default the deploy script also refreshes Z:\LLM\model-chat-app before restarting the selected
service, so changes in scripts/transformers_plugin_server.py are applied to the mounted /app
directory. Use -NoSyncApp only for a fast container restart when the app files are already current.
Ports:
- translation:
8010 - audio:
8020 - video:
8030 - image:
8040
The translation service exposes /health, /v1/models, and /v1/chat/completions.
The audio service exposes /health, /v1/models, and /v1/audio/transcriptions; in the UI,
select plugin Звук, choose an audio file, then press Отправить.
The video service exposes /health, /v1/models, and /v1/vision/analyze; in the UI,
select plugin Видео, choose an image, enter the question in the prompt box, then press Отправить.
The image service exposes /health, /v1/models, /v1/images/generations, /v1/images/edits,
and async job endpoints under /v1/images/jobs. In the UI, select plugin Фото, choose
an image model, choose generate or edit / inpaint, enter the prompt, then press Отправить.
By default LOAD_ON_START=1 and BACKGROUND_LOAD_ON_START=1 for image generation, so the service
opens HTTP quickly and loads the SDXL base model in the background. The first request can show
loading_model for several minutes. Switching between SDXL base and inpaint may unload the other
pipeline to keep VRAM available.
Stop the large text vLLM container before loading another large model if VRAM is tight.
The UI can select Qwen Image Edit, but the image endpoint must actually serve
qwen-image-edit; otherwise /api/image/submit rejects the job instead of silently using SDXL.
To test Qwen Image Edit as a single-heavy-model experiment, restart the image service with the Qwen
env file:
docker -H ssh://docker-gpu compose `
--env-file core/deploy/docker-gpu/transformers/image.qwen-edit.env.example `
-f core/deploy/docker-gpu/transformers/image.compose.yaml up -d
Restore the verified SDXL service with:
powershell -ExecutionPolicy Bypass -File scripts\deploy_transformers_service.ps1 -Plugin image
Verified Qwen Image Edit notes from 2026-06-20:
- model files are present under
Z:\LLM\models\image\qwen-image-edit; diffusersdetectsQwenImageEditPipelineand the service can expose/v1/modelsasqwen-image-edit;- pipeline load completed in about
31 swith CPU offload on RTX 4090; - a 512x512 edit job with
1inference step did not finish within1800 s, so Qwen Image Edit is not practical on the current 24 GB GPU profile without a quantized/optimized runtime or a larger GPU; - keep SDXL as the default verified image service for now.
Smoke test image jobs through the Model Chat proxy:
python scripts\smoke_image_jobs.py --operation generate --steps 4
python scripts\smoke_image_jobs.py --operation edit --steps 4
python scripts\smoke_image_jobs.py --operation edit --steps 40 --cancel --cancel-after 1
python scripts\smoke_image_jobs.py --model-id qwen-image-edit --model-mode qwen-image-edit --model qwen-image-edit --operation edit --steps 4
Audio and video smoke checks generate small synthetic inputs locally and send them through the Model Chat UI proxy:
powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile audio
python scripts\smoke_audio_transcription.py
powershell -ExecutionPolicy Bypass -File scripts\switch_gpu_profile.ps1 -Profile video
python scripts\smoke_video_analysis.py
When the matching service is intentionally stopped, use --allow-unavailable for a code-path check
that does not fail the local validation run.
Verified on 2026-06-20 with RTX 4090:
- generate, SDXL base, 512x512, 4 steps: completed in
2194 msafter warmup; - edit / inpaint, SDXL inpaint, 512x512, 4 steps: completed in
180649 msincluding first inpaint model load; - generate, SDXL base, 512x512, 1 step: completed in
86866 msafter image service restart and warmup; - generated artifacts are served through
/generated-images/<date>/<file>.png.
Direct artifact check uses GET; this minimal HTTP server does not implement HEAD for generated files.
Model Downloads On GPU Host
Use this when model files should be written directly to Z:\LLM\models on the GPU host:
For large Hugging Face shard files, prefer the explicit range downloader. It writes final files directly
to /models/..., resumes by local file size, and avoids stale .cache/huggingface/download/*.incomplete
files left by interrupted Xet/snapshot downloads.
powershell -ExecutionPolicy Bypass -File scripts\download_hf_range_gpu.ps1 `
-CardId qwen2_5-vl-7b-instruct `
-LocalDir /models/video/qwen2.5-vl-7b-instruct `
-AllowFile "model-00001-of-00005.safetensors,model-00002-of-00005.safetensors,model-00003-of-00005.safetensors,model-00004-of-00005.safetensors,model-00005-of-00005.safetensors" `
-Detached `
-ContainerName llm-hf-range-qwen-vl `
-ChunkSize 256mb
The registry/snapshot downloader is still useful for dry-run planning and small metadata files:
powershell -ExecutionPolicy Bypass -File scripts\download_missing_hf_models_gpu.ps1 -DryRun
docker -H ssh://docker-gpu rm -f llm-model-download
powershell -ExecutionPolicy Bypass -File scripts\download_missing_hf_models_gpu.ps1 `
-Detached `
-ContainerName llm-model-download
Monitor:
docker -H ssh://docker-gpu ps
docker -H ssh://docker-gpu logs --tail 100 llm-model-download
Invoke-RestMethod http://192.168.220.91:8765/api/model-services
Notes
- UI is served by
scripts/model_chat_server.py. - Inference endpoint defaults to
http://docker-gpu.cin.su:8000. - Use
vllm/vllm-openai:v0.10.2for the UI image because it already contains Python and is verified on this host. vllm/vllm-openai:latestpasses CUDA after driver595.97, but vLLM0.23.0currently fails on Docker Desktop/WSL withUVA is not available.