Back to the expert workbench

THE NEXT EXPERIMENT

From expert edits to a better Gemma.

The workbench collects examples. Approving a response does not start training. We first establish whether the original model produces useful drafts, then teach it with expert corrections.

1. Collect a small, strong dataset

Start with 10–15 baseline tasks. Build toward 50–100 carefully edited responses for a pilot. Correct the entire response: the Specific Aim and its missing-information notes. Record scientific errors, usefulness, and editing time.

2. Keep projects separate

Choose training, validation, or held-out evaluation for each project. Keep related aims and excerpts together. Never train on held-out answers. The first approved example fixes that project’s assignment.

3. Export approved examples

The training export contains only approved responses with permission. It includes separate files for each nonempty split, plus a manifest. Check near duplicates, sources, and the exact input/answer pairs before training.

4. Run a small QLoRA experiment

The hosted workers now serve Gemma 4 12B through MLX, sharing the existing 4-bit checkpoint. Literature adaptation is separate from this planned grant-writing supervised stage. Use a new run directory, the native Gemma 4 chat template and assistant-only loss. Start with a short compatibility run. Pause both workers and release loaded models before training.

The measured hardware is an Apple M4 with 24 GiB RAM. The paper experiment completed 152 training steps on 15 primary papers; grant-writing SFT has not been run. Memory and speed at longer grant-writing context lengths still need measurement. Reuse the pinned MLX checkpoint; no second Ollama model download is needed.

5. Compare fairly

Compare the untuned model and adapter in the same runtime with identical evidence and settings. Blind the model labels. Prioritize scientific fidelity and reduced editing time. A smoother draft that invents more claims fails.

6. Release the winning model

Keep the baseline available. The existing MLX workers can serve an approved adapter directly; update the supported run and its hashes after evaluation. Gemma 3 adapters cannot be transferred to Gemma 4. Preference tuning can wait until supervised tuning has been evaluated.

Implementation references