Experiments Attempted But ..
A catalog of 34 experiments from this project's lab notebook — wins, losses, null results, and a few outright breaking points — each with its own result page.
We Can Not Only Learn Activations, But Also Transfer Them
One shared activation, five numbers, learned on ImageNet-1K — then frozen and dropped into a 25x smaller Vision Transformer, where it still beats ten fixed activations.
One Learnable Activation, Shared Across the Whole Network — Outperforms Fixed GELU at ImageNet-1K Scale
One learnable nonlinearity, shared by every layer, tested head-to-head against a standard fixed activation on real ImageNet-1K.
Thinking beyond localized activations
Every network you’ve trained almost certainly leans on the same handful of
activation functions — ReLU, GELU, SiLU, tanh. They differ in shape, but they
share one quiet assumption: the nonlinearity itself is fixed. You pick one
function, wide enough to cover the whole range of inputs a neuron will ever
see, bolt it on after every layer, and let ...
வேதாந்தக் குறிப்புகள் — குருமுகமாகக் கேட்டறிந்தவை
गुरुर्ब्रह्मा गुरुर्विष्णुः गुरुर्देवो महेश्वरः ।
गुरुः साक्षात् परं ब्रह्म तस्मै श्रीगुरवे नमः ॥
guru-r-brahmā guru-r-viṣṇuḥ guru-r-devo maheśvaraḥ,
guruḥ sākṣāt paraṁ brahma tasmai śrī-gurave namaḥ.
குரு பிரம்மா; குரு விஷ்ணு; குரு தேவனாகிய மகேஸ்வரன்; குருவே சாட்சாத் பரப்பிரம்மம்; அத்தகைய குருவுக்கு நமஸ்காரம்.
ஸ்ரீ சங்கரானந்தா
இக்கட...
More Is Different
When we scale a neural network to billions of parameters, new capabilities emerge. But we still lack a complete explanation of why certain capabilities emerge or how they arise. Well, this is not unique to neural networks!
I recently came across a beautiful essay, “More Is Different,” by Philip W. Anderson. Anderson argued that as systems becom...
Open Agents for Building Interactive Simulations
Recently, I built a set of interactive simulations for teaching math and science concepts to kids and students. However, I have used Claude to generate such simulations. I wondered, can we build such applets using open-weight models? It turns out that we CAN. We do not need very big models with hundreds of billions of parameters. With an agentic...
Building Interactive Simulations for My Daughter
My 6-year-old daughter asked me about the phases of the moon. I tried explaining it using pen and paper, hand-waving in the air — she got a bit of it. Like every kid, she watches space cartoons and videos on YouTube. Suddenly I thought of building an interactive simulation like in PhET. I searched PhET, and I didn’t find any app for explaining m...
37 post articles, 5 pages.