Radius looks like a single Greek letter until someone proposes enlarging it mid-meeting.

In PeTeR’s framing, post-training robustification asks a pre-trained circuit to remain reliable against the worst case inside an ε\varepsilon-Circuit-Wasserstein ball. Choose ε\varepsilon too small and you have cosmetics. Choose it too large and you are optimizing against a neighborhood so wide that “robust” becomes a mood rather than a guarantee.

Why the metric matters

Because CW\mathsf{CW} upper-bounds classical Wasserstein, an ε\varepsilon-CW\mathsf{CW} ball is a stricter neighborhood than the corresponding Big Wasserstein ball. That is not a bug; it is why the community can finish a computation before lunch. Members who insist on “the true ball” are usually inviting Big Wasserstein back into the room, along with its #P\#\mathsf{P}-hard manners.

Formally, the uncertainty set is

{Q:CW(P^,Q)ε},\{ Q : \mathsf{CW}(\hat{P}, Q) \le \varepsilon \},

and the robustification objective hedges against the worst QQ inside it.

A practical checklist

Before you publish an ε\varepsilon, answer:

  • Compatible structure, or are we about to pay for a compatibility transform?
  • Is the adversary parameterized as a circuit we can actually optimize against?
  • Would this radius still make sense if OpenCW announced the same number with none of the proofs?

If those answers are available without theatrics, the radius is probably fine.

Consensus is not the absence of disagreement. It is disagreement that has been given a constraint set.

Next issue records minutes from a meeting in which snacks briefly outranked subgradients.

← Back to posts