GRZO
Group-Relative Zeroth-Order Optimization for Large Language Model Fine-Tuning
One perturbation per example, not one per batch.
University of California, Santa Barbara
Overview
One perturbation per example, not one per batch.
Zeroth-order (ZO) fine-tuning trains an LLM with forward passes only, so it fits in inference-level memory. Its weak spot is the gradient estimate: MeZO draws one random perturbation for the whole mini-batch, so every step follows a single, very noisy direction.
GRZO gives every example its own pseudo-independent perturbation and combines the per-example loss differences with a group-relative normalization. One step now carries B directions instead of one — with the same two forward passes, and peak memory still at the inference floor.
On Llama3-8B, GRZO beats MeZO by +3.0 average accuracy while staying within 0.5% of the forward-only memory floor. It is a drop-in replacement for the MeZO core: inside sparse, low-rank and quantized ZO methods it lifts them by +4.9 on average.
Method
Turn the batch dimension into the variance-reduction lever.
Earlier fixes either shrink the effective parameter dimension (and lose full-parameter expressivity) or enrich the estimator with extra forward passes or persistent state. GRZO keeps MeZO's envelope — two forward passes, inference memory — and gets its extra directions from the examples it already processes.
Per-example perturbations
One shared base perturbation U per layer, modulated by per-example random sign vectors (a Flipout-style factorization). B pseudo-independent directions, no B weight copies.
Group-relative weights
Each example's two-sided loss difference δi is normalized by the within-batch standard deviation, as in GRPO advantages — scale-free and robust to loss-magnitude drift.
Same budget
Perturbations are built transiently inside each layer's forward hook and regenerated from seeds for the update, so the B directions cost +0.08 GB on Llama3-8B.
\(r_i, s_i\) are per-example Rademacher sign vectors, \(\delta_i = \ell_i^{+}-\ell_i^{-}\) the two-sided loss difference of example \(i\), \(s\) the within-batch standard deviation of the \(\delta_i\), and \(z_i=\mathrm{vec}(\Delta W_i)\).
Why it works
B directions buy a Beff-fold variance cut.
GRZO is directionally unbiased, and its variance relative to MeZO is governed by how aligned the per-example gradients are. With c the mean pairwise cosine between them:
So GRZO reaches an \(\epsilon\)-stationary point in about \(B_{\text{eff}}\times\) fewer steps. On a Llama3-8B decoder block the measured alignment is \(c = 0.83\), which gives \(B_{\text{eff}} \approx 13\) at \(B = 16\). Try it below.
Results
Best zeroth-order accuracy, at inference memory.
Full-parameter fine-tuning for 20k steps at batch size 16 on SuperGLUE classification and QA. GRZO is the strongest ZO method on Llama3-8B (+3.0 average over MeZO, +7.2 on RTE) and on OPT-13B, narrowing the gap to first-order Adam without its memory.
Cost
What the extra B−1 directions cost.
Production profile on Llama3-8B (RTE, fp16, B = 16, 4×A100). Memory: +0.08 GB, or 0.5% above the forward-only floor — GRZO does not use less memory than an in-place MeZO, it gets B directions at the same footprint. Time: +23% per step, from fusing the per-example perturbation into the forward pass.
Drop-in
Swap the MeZO core, lift every variant.
Sparse, low-rank and quantized ZO methods all sit on top of a MeZO-style estimator. Replacing that core with GRZO — and changing nothing else — improves the paired method on 14 of 15 task–variant combinations, by +4.9 on average (Llama3-8B).
Takeaway. The B per-example directions are orthogonal to how a variant restricts or compresses the update, so the gains stack.
FAQ
Questions people ask
What problem does GRZO solve?
It reduces the variance of zeroth-order gradient estimation when fine-tuning large language models, which is the main reason ZO fine-tuning underperforms backpropagation.
How is GRZO different from MeZO?
MeZO uses one random perturbation shared by the whole mini-batch, so each step yields a single gradient direction. GRZO assigns a pseudo-independent perturbation to every example in the batch and combines the per-example losses with group-relative normalization, so one step yields as many directions as there are examples — without any extra forward passes.
Does GRZO cost more compute or memory than MeZO?
Memory, no: the number of forward passes is unchanged and peak memory stays at the inference floor (17.82 GB on Llama3-8B, 0.5% above the forward-only floor, essentially level with MeZO). Time, slightly: GRZO carries a 23% per-step time premium, but it converges fastest in wall-clock time, so it reaches a given loss sooner.
How much better is it in practice?
Average accuracy on Llama3-8B improves by +3.0 over MeZO, and on RTE it reaches 81.6% against MeZO's 74.4%. Used as a drop-in replacement for the MeZO core inside sparse, low-rank, and quantized ZO methods, it improves them by +4.9 on average. Evaluated on RoBERTa-large, Llama3-8B, and OPT-13B.
Is there a theoretical guarantee?
Yes. GRZO is shown to be directionally unbiased, its variance shrinks proportionally to the batch size, and it admits a tighter nonconvex convergence bound than MeZO.
When should I use GRZO instead of backpropagation?
When memory is the binding constraint — fine-tuning a model that does not fit in memory with activations stored, or on hardware without enough capacity for a backward pass. GRZO keeps ZO's inference-level memory while closing part of the accuracy gap.
Citation
BibTeX
Abstract
Zeroth-order (ZO) optimization is a memory-efficient alternative to backpropagation for fine-tuning large language models, but its deployment is limited by the high variance of gradient estimation. We propose GRZO, a Group-Relative Zeroth-Order optimizer that draws one pseudo-independent perturbation per mini-batch example and aggregates the per-example losses through group-relative normalization, raising the effective gradient-direction count from one to the batch size at no additional forward cost while preserving inference-level memory. We prove that GRZO is directionally unbiased with variance shrinking proportionally to the batch size, yielding a tighter nonconvex convergence bound than MeZO. Across RoBERTa-large, Llama3-8B, and OPT-13B over multiple tasks, GRZO improves average accuracy on Llama3-8B by +3.0 over MeZO while staying within 0.5% of the forward-only inference memory floor; as a drop-in replacement for the MeZO core, it lifts sparse, low-rank, and quantized ZO variants by +4.9 on average.
@inproceedings{tan2026grzo,
title = {GRZO: Group-Relative Zeroth-Order Optimization for Large Language Model Fine-Tuning},
author = {Tan, Liyan and Zhao, Yequan and Yang, Yifan and Zhang, Ruijie and Yu, Xinling and Zhang, Zheng},
booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026},
year = {2026}
}