Interactive figure explorer

Towards Transferable Activation Functions

Explore the paper interactively. Every figure and table is live, with filters for activation, dataset, configuration, and seed. FAct is a single five-parameter nonlinearity shared across all MLP blocks, learned jointly with a depth-6 ViT on ImageNet-1K, then frozen and transferred—without source weights—to an independently initialized depth-2 ViT. Drag to pan, scroll to zoom, and toggle individual arms or seeds.

Figure 1 · Section 3.2–3.3Activation Parameterization and Sharing

FAct is five numbers. Move them and watch the whole network's nonlinearity change. The shaded strip is where the source model's pre-activations actually live — only the shape inside it does any work.

The shared curve φ(z)

Scroll to zoom, drag to pan, double-click to reset.

Fourier coefficients (K = 2)

φ(z) = a₀ + Σₖ aₖcos(kωz) + bₖsin(kωz)

Activation parameters vs. network size

Depth-6 ViT, 3.05M weights. Log scale.

Figure 3 · Section 4.2The Shared Nonlinearity Starts Near GELU and Evolves Away From It

Scrub through the 100 ImageNet-1K epochs. The curve is initialized from a two-harmonic Fourier fit of GELU and then leaves it.

100

Learned curve at the selected epoch

Dotted = GELU-fit initialization. Grey = fixed activations.

Coefficient trajectories

All five coefficients across training. Vertical line marks the selected epoch.

Where the pre-activations are

Pooled over 512 ImageNet images and all six blocks of the source ViT-d6.

Table 1 & Figure 4 · Section 4.2Accuracy and Optimization on ImageNet-1K

Depth-6 ViT, 3.05M parameters, 100 epochs. Learnable FAct and GELU are five-seed means with a ±1 s.d. band; the other arms are single runs.

Validation top-1 against epoch

Scroll to zoom the epoch axis, drag to pan, double-click to reset.

Seed-matched gap: Learnable FAct − GELU

Thin lines are the five seed-matched pairs; the thick line is their mean.

Tables 2–3 · Section 4.3–4.4Frozen Transfer to an Independently Initialized ViT-d2

Only the five coefficients cross over — no source weights, no further activation updates. Every dot is one seed.

Target test accuracy by activation

Bar = mean, whisker = ±1 s.d., dots = individual seeds. Hover for the seed values.

Table view

Target training trajectories

Seed-mean validation accuracy with a ±1 s.d. band. Toggle arms below.

Table 3 · Section 4.4Robustness to Activation-Specific Hyperparameter Tuning

Each activation gets its own learning rate and weight decay, selected on a held-out split. The lines show what tuning does to every arm on the same axis.

Common configuration → activation-specific configuration

Hover a line for the selected (lr, wd) and the change in accuracy.

Table 4 · Section 4.5Comparison Across Activation Choices

Eleven activation choices × four target datasets. Colour ranks each column on its own scale — the legend names which end is strongest — and the spread between the strongest and weakest fixed activation is printed above each column.

Activation zoo

Hover a cell for the mean, s.d., seed count and rank.

Table view

Appendix BDetailed Transfer Statistics

Seed-matched margins of Frozen FAct over each fixed activation, in percentage points. A margin whose interval clears zero is a win on every shared seed.

Frozen FAct − fixed activation (pp)

Dot = mean paired difference, bar = ±1 s.d. of the paired differences. Hover for the paired t, Holm-adjusted p, and the win count.

Table view

Appendix D · Section 4.6Is Transferability Specific to the Fourier Parameterization?

A globally shared Padé activation, learned the same way, is a different function in a different basis. Over the range the target network actually uses, it transfers comparably.

The two frozen curves

Shaded histogram: the target network's own pre-activations on this dataset.

Transfer accuracy, common configuration

Six seeds per cell. All three arms were re-run on one GPU so that no comparison crosses hosts, so these differ slightly from the Frozen-transfer panel. Hover for the paired comparison against Frozen PAU.

Appendix EKnowledge Distillation

A ViT-L/16 teacher distilled into the same 3.05M-parameter student. Distillation and a learned activation add up rather than substitute for each other.

Validation top-1 against epoch

Solid = with distillation, dashed = without. Seed 1.

Does the teacher reshape the curve?

Each coefficient, with and without distillation, across training.

Appendix GFine-Tuning the Source Network on Food-101

The fourth transfer mode: carry the whole trained source network over, not just the activation. The learning-rate sweep below is the search that picked 1×10⁻⁴.

Learning-rate search

Seed 0. Every transplanted tensor trains; the marker is the selected rate.

Margin over GELU at each learning rate

Percentage points. The advantage is largest where the rate is smallest.