1.1 KiB
1.1 KiB
First Model Selection
Selected Model
Qwen/Qwen3-4B-Instruct-2507
Local served name: qwen3-4b-instruct.
Why This Model
- Apache-2.0 license.
- 4B parameters, so it is a practical first GPU deployment target.
- Supports text generation, chat, coding, multilingual tasks and tool-use scenarios.
- Official model page includes vLLM usage.
- Native context length is 262,144 tokens, but the first deployment uses 32,768 tokens to reduce OOM risk until GPU VRAM is confirmed.
Alternatives
Qwen/Qwen3-8B: stronger base candidate, but requires more VRAM.Qwen/Qwen3.5-4B: newer and long-context oriented, but its model card currently recommends newer/mainline serving frameworks, so it is better as a second experiment after the first stable deployment.- Mistral/DeepSeek coder models: useful future candidates for coding-specific comparison, but not the first default for Russian + 1C + general assistant coverage.