Explore the paper interactively. Every figure and table is live, with filters for activation, dataset, configuration, and seed. FAct is a single five-parameter nonlinearity shared across all MLP blocks, learned jointly with a depth-6 ViT on ImageNet-1K, then frozen and transferred—without source weights—to an independently initialized depth-2 ViT. Drag to pan, scroll to zoom, and toggle individual arms or seeds.
FAct is five numbers. Move them and watch the whole network's nonlinearity change. The shaded strip is where the source model's pre-activations actually live — only the shape inside it does any work.
Scroll to zoom, drag to pan, double-click to reset.
φ(z) = a₀ + Σₖ aₖcos(kωz) + bₖsin(kωz)
Depth-6 ViT, 3.05M weights. Log scale.
Scrub through the 100 ImageNet-1K epochs. The curve is initialized from a two-harmonic Fourier fit of GELU and then leaves it.
Dotted = GELU-fit initialization. Grey = fixed activations.
All five coefficients across training. Vertical line marks the selected epoch.
Pooled over 512 ImageNet images and all six blocks of the source ViT-d6.
Depth-6 ViT, 3.05M parameters, 100 epochs. Learnable FAct and GELU are five-seed means with a ±1 s.d. band; the other arms are single runs.
Scroll to zoom the epoch axis, drag to pan, double-click to reset.
Thin lines are the five seed-matched pairs; the thick line is their mean.
Only the five coefficients cross over — no source weights, no further activation updates. Every dot is one seed.
Bar = mean, whisker = ±1 s.d., dots = individual seeds. Hover for the seed values.
Seed-mean validation accuracy with a ±1 s.d. band. Toggle arms below.
Each activation gets its own learning rate and weight decay, selected on a held-out split. The lines show what tuning does to every arm on the same axis.
Hover a line for the selected (lr, wd) and the change in accuracy.
Eleven activation choices × four target datasets. Colour ranks each column on its own scale — the legend names which end is strongest — and the spread between the strongest and weakest fixed activation is printed above each column.
Hover a cell for the mean, s.d., seed count and rank.
Seed-matched margins of Frozen FAct over each fixed activation, in percentage points. A margin whose interval clears zero is a win on every shared seed.
Dot = mean paired difference, bar = ±1 s.d. of the paired differences. Hover for the paired t, Holm-adjusted p, and the win count.
A globally shared Padé activation, learned the same way, is a different function in a different basis. Over the range the target network actually uses, it transfers comparably.
Shaded histogram: the target network's own pre-activations on this dataset.
Six seeds per cell. All three arms were re-run on one GPU so that no comparison crosses hosts, so these differ slightly from the Frozen-transfer panel. Hover for the paired comparison against Frozen PAU.
A ViT-L/16 teacher distilled into the same 3.05M-parameter student. Distillation and a learned activation add up rather than substitute for each other.
Solid = with distillation, dashed = without. Seed 1.
Each coefficient, with and without distillation, across training.
The fourth transfer mode: carry the whole trained source network over, not just the activation. The learning-rate sweep below is the search that picked 1×10⁻⁴.
Seed 0. Every transplanted tensor trains; the marker is the selected rate.
Percentage points. The advantage is largest where the rate is smallest.