CSE 151B Competition
Reasoned through math problems with Qwen3-4B.
Course competition · Spring 2026

Private score
69.9
Weights
Qwen3-4B, no FT
Private set
943 questions
Problem
The private set is 943 competition-math questions (300 MCQ, 643 free-response). Fine-tuning a 4B model on Hub chain-of-thought data looked like the obvious path, but LoRA adapters lost to the untouched base weights on the public dev set.
Approach
Kept Qwen3-4B-Thinking frozen and served it with vLLM plus bitsandbytes INT8. MCQ uses thinking mode, letter extraction, and a short greedy finalizer on the noisiest 25% of each batch. Free-response is a single long decode plus boxed-answer canonicalization. A second submission sampled three FRQ completions and majority-voted (68.9); the default single pass scored higher.
In short
UC San Diego CSE 151B competition pipeline. Unmodified Qwen/Qwen3-4B-Thinking-2507 served with vLLM INT8. The best private run (69.9) is a single free-response pass with thinking-mode MCQ, a greedy finalizer on noisy batches, and judger-ready boxed answers. A second run with 3-sample FRQ self-consistency scored 68.9. LoRA adapters were trained and documented; they lost to the base model, so they were not submitted.
Stack
- Python
- vLLM
- Qwen3
- LLM Inference
- Bitsandbytes