Back to projects

CSE 151B Competition

Reasoned through math problems with Qwen3-4B.

Course competition · Spring 2026

CSE 151B competition card: fine-tuning Qwen3-4B

Problem

The private set is 943 competition-math questions (300 MCQ, 643 free-response). Fine-tuning a 4B model on Hub chain-of-thought data looked like the obvious path, but LoRA adapters lost to the untouched base weights on the public dev set.

Approach

Kept Qwen3-4B-Thinking frozen and served it with vLLM plus bitsandbytes INT8. MCQ uses thinking mode, letter extraction, and a short greedy finalizer on the noisiest 25% of each batch. Free-response is a single long decode plus boxed-answer canonicalization. A second submission sampled three FRQ completions and majority-voted (68.9); the default single pass scored higher.

In short

UC San Diego CSE 151B competition pipeline. Unmodified Qwen/Qwen3-4B-Thinking-2507 served with vLLM INT8. The best private run (69.9) is a single free-response pass with thinking-mode MCQ, a greedy finalizer on noisy batches, and judger-ready boxed answers. A second run with 3-sample FRQ self-consistency scored 68.9. LoRA adapters were trained and documented; they lost to the base model, so they were not submitted.

Stack

  • Python
  • vLLM
  • Qwen3
  • LLM Inference
  • Bitsandbytes