GRZO
Findings of EMNLP 2026

GRZO

Group-Relative Zeroth-Order Optimization for Large Language Model Fine-Tuning

One perturbation per example, not one per batch.

University of California, Santa Barbara

+3.0average accuracy over MeZO on Llama3-8B
16gradient directions per step at B = 16 (MeZO: 1)
0.5%above the forward-only inference memory floor
+4.9average lift for sparse, low-rank and quantized ZO variants

Overview

One perturbation per example, not one per batch.

Zeroth-order (ZO) fine-tuning trains an LLM with forward passes only, so it fits in inference-level memory. Its weak spot is the gradient estimate: MeZO draws one random perturbation for the whole mini-batch, so every step follows a single, very noisy direction.

GRZO gives every example its own pseudo-independent perturbation and combines the per-example loss differences with a group-relative normalization. One step now carries B directions instead of one — with the same two forward passes, and peak memory still at the inference floor.

TL;DR

On Llama3-8B, GRZO beats MeZO by +3.0 average accuracy while staying within 0.5% of the forward-only memory floor. It is a drop-in replacement for the MeZO core: inside sparse, low-rank and quantized ZO methods it lifts them by +4.9 on average.

Side-by-side pipeline comparison of MeZO and GRZO
MeZO vs. GRZO. MeZO (left) shares one perturbation across the mini-batch: one gradient direction per step. GRZO (right) builds pseudo-independent per-example perturbations and applies group-relative normalization: B directions per step, much lower variance, the same forward budget.

Method

Turn the batch dimension into the variance-reduction lever.

Earlier fixes either shrink the effective parameter dimension (and lose full-parameter expressivity) or enrich the estimator with extra forward passes or persistent state. GRZO keeps MeZO's envelope — two forward passes, inference memory — and gets its extra directions from the examples it already processes.

1

Per-example perturbations

One shared base perturbation U per layer, modulated by per-example random sign vectors (a Flipout-style factorization). B pseudo-independent directions, no B weight copies.

2

Group-relative weights

Each example's two-sided loss difference δi is normalized by the within-batch standard deviation, as in GRPO advantages — scale-free and robust to loss-magnitude drift.

3

Same budget

Perturbations are built transiently inside each layer's forward hook and regenerated from seeds for the update, so the B directions cost +0.08 GB on Llama3-8B.

$$\Delta W_i = U \odot (r_i s_i^{\top}), \qquad a_i = \frac{\delta_i}{s+\epsilon}, \qquad \hat g = \frac{1}{2\sigma B}\sum_{i=1}^{B} a_i\, z_i$$

\(r_i, s_i\) are per-example Rademacher sign vectors, \(\delta_i = \ell_i^{+}-\ell_i^{-}\) the two-sided loss difference of example \(i\), \(s\) the within-batch standard deviation of the \(\delta_i\), and \(z_i=\mathrm{vec}(\Delta W_i)\).

Why it works

B directions buy a Beff-fold variance cut.

GRZO is directionally unbiased, and its variance relative to MeZO is governed by how aligned the per-example gradients are. With c the mean pairwise cosine between them:

$$\frac{\mathrm{Var}(\hat g_{\text{MeZO}})}{\mathrm{Var}(\hat g_{\text{GRZO}})} \;\approx\; B_{\text{eff}} \;=\; cB + (1-c)$$

So GRZO reaches an \(\epsilon\)-stationary point in about \(B_{\text{eff}}\times\) fewer steps. On a Llama3-8B decoder block the measured alignment is \(c = 0.83\), which gives \(B_{\text{eff}} \approx 13\) at \(B = 16\). Try it below.

Watch the estimate settle Toy simulation

Each arrow is one step's gradient estimate, drawn at its true angle to the exact batch gradient (dashed, pointing up). 32-dimensional Gaussian toy — an illustration of the theorem, not experimental data.

MeZO1 direction / step
GRZO16 directions / step
16
0.83
MeZO mean cosine–
GRZO mean cosine–
Predicted BeffcB + (1 − c)–

Measured on Llama3-8B

Cosine similarity between a single ZO estimate and the exact backpropagation gradient on one Llama3-8B decoder block (218M parameters), mean over 3,000 paired trials per B. In 218M dimensions every cosine is tiny; what matters is that GRZO's grows with B while MeZO's stays flat.

GRZOMeZO
At B = 16×4.6
At B = 32×6.5

GRZO’s cosine over MeZO’s. MeZO stays near 5.3×10−5 at every B: one shared direction, however many examples.

Results

Best zeroth-order accuracy, at inference memory.

Full-parameter fine-tuning for 20k steps at batch size 16 on SuperGLUE classification and QA. GRZO is the strongest ZO method on Llama3-8B (+3.0 average over MeZO, +7.2 on RTE) and on OPT-13B, narrowing the gap to first-order Adam without its memory.

Accuracy by task

Accuracy (%) for classification, F1 for SQuAD and DROP. Each column is a task: the line rises from MeZO to GRZO, and the black dash is first-order Adam, which needs full backpropagation memory. The top row is GRZO’s gain over MeZO.

MeZOFZOOGRZO (ours)Adam (first-order)
Show the full table (incl. LoRA)

Convergence: hover to compare

Training loss as in the paper (Llama3-8B: MultiRC, RTE; OPT-13B: DROP, SQuAD), plotted against steps or wall-clock time. Move the cursor: the dashed line is GRZO’s loss at that point, and the rings mark where each baseline first gets that low.

GRZOFZOOMeZO

Curves digitized from the paper’s convergence figure (smoothed as plotted there; FZOO on RTE smoothed a little more), clipped to the same window; MeZO diverges on DROP. GRZO’s 23% per-step premium is already included in the wall-clock panels.

Cost

What the extra B−1 directions cost.

Production profile on Llama3-8B (RTE, fp16, B = 16, 4×A100). Memory: +0.08 GB, or 0.5% above the forward-only floor — GRZO does not use less memory than an in-place MeZO, it gets B directions at the same footprint. Time: +23% per step, from fusing the per-example perturbation into the forward pass.

Peak memory and per-step time

Each ZO variant paired with its GRZO version. Every bar starts at zero; the vertical line marks the 17.74 GB forward-only inference peak.

Peak GPU memoryGB

Per-step timems

Drop-in

Swap the MeZO core, lift every variant.

Sparse, low-rank and quantized ZO methods all sit on top of a MeZO-style estimator. Replacing that core with GRZO — and changing nothing else — improves the paired method on 14 of 15 task–variant combinations, by +4.9 on average (Llama3-8B).

Variant before → after

Each column rises from the published variant (orange dot) to the same variant with the GRZO core swapped in (blue dot). The hollow ring is vanilla GRZO for reference; the top row is the gain.

Takeaway. The B per-example directions are orthogonal to how a variant restricts or compresses the update, so the gains stack.

FAQ

Questions people ask

What problem does GRZO solve?

It reduces the variance of zeroth-order gradient estimation when fine-tuning large language models, which is the main reason ZO fine-tuning underperforms backpropagation.

How is GRZO different from MeZO?

MeZO uses one random perturbation shared by the whole mini-batch, so each step yields a single gradient direction. GRZO assigns a pseudo-independent perturbation to every example in the batch and combines the per-example losses with group-relative normalization, so one step yields as many directions as there are examples — without any extra forward passes.

Does GRZO cost more compute or memory than MeZO?

Memory, no: the number of forward passes is unchanged and peak memory stays at the inference floor (17.82 GB on Llama3-8B, 0.5% above the forward-only floor, essentially level with MeZO). Time, slightly: GRZO carries a 23% per-step time premium, but it converges fastest in wall-clock time, so it reaches a given loss sooner.

How much better is it in practice?

Average accuracy on Llama3-8B improves by +3.0 over MeZO, and on RTE it reaches 81.6% against MeZO's 74.4%. Used as a drop-in replacement for the MeZO core inside sparse, low-rank, and quantized ZO methods, it improves them by +4.9 on average. Evaluated on RoBERTa-large, Llama3-8B, and OPT-13B.

Is there a theoretical guarantee?

Yes. GRZO is shown to be directionally unbiased, its variance shrinks proportionally to the batch size, and it admits a tighter nonconvex convergence bound than MeZO.

When should I use GRZO instead of backpropagation?

When memory is the binding constraint — fine-tuning a model that does not fit in memory with activations stored, or on hardware without enough capacity for a backward pass. GRZO keeps ZO's inference-level memory while closing part of the accuracy gap.

Citation

BibTeX

Abstract

Zeroth-order (ZO) optimization is a memory-efficient alternative to backpropagation for fine-tuning large language models, but its deployment is limited by the high variance of gradient estimation. We propose GRZO, a Group-Relative Zeroth-Order optimizer that draws one pseudo-independent perturbation per mini-batch example and aggregates the per-example losses through group-relative normalization, raising the effective gradient-direction count from one to the batch size at no additional forward cost while preserving inference-level memory. We prove that GRZO is directionally unbiased with variance shrinking proportionally to the batch size, yielding a tighter nonconvex convergence bound than MeZO. Across RoBERTa-large, Llama3-8B, and OPT-13B over multiple tasks, GRZO improves average accuracy on Llama3-8B by +3.0 over MeZO while staying within 0.5% of the forward-only inference memory floor; as a drop-in replacement for the MeZO core, it lifts sparse, low-rank, and quantized ZO variants by +4.9 on average.

@inproceedings{tan2026grzo,
  title     = {GRZO: Group-Relative Zeroth-Order Optimization for Large Language Model Fine-Tuning},
  author    = {Tan, Liyan and Zhao, Yequan and Yang, Yifan and Zhang, Ruijie and Yu, Xinling and Zhang, Zheng},
  booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026},
  year      = {2026}
}