Home

Masked Attentions in Transformer Architectures

Introduction Masked attention is typically used in the Decoder part of (vanilla) transformer architecture to prevent the model from looking at future tokens. This left me with the impression that the models trained using Masked Language Modelling (MLM) objective use Masked Attention. However, masked language models like BERT (Bidirectional En...

Read more

Running Jupyter Notebook from a Remote Server

Introduction If you are training a deep learning model or fine-tuning LLMs (Large Language Models), then at some point in time you need to connect with a remote machine that has required amount of computing power (about 80GB or 320 GB of GPU Memory). Data scientists often use Jupyter Notebooks for experimentation. Jupyter Notebook is designed t...

Read more

Pytorch for Deep Learning

Theory is not enough, you must apply. The objective of this workshop is to give you hands on experience in building models using the pytorch's core component called Tensor . Beleive me, everything you are gonna build is simply stacking or connecting copies of this single core component!. It make sense as all the models takes in Tensors (data) a...

Read more

கணிப்போம் வா

"கண்ணா, வெளியில விளையாடப் போறயா, குடை எடுத்துட்டு போப்பா, மழை வர 60% வாய்ப்பிருக்குனு (Chance) போட்ருக்குப்பா ", என்றாள் திண்ணையில் ஸ்மார்ட் போனை ஸ்க்ரால் செய்தபடி அமர்ந்திருந்த பாட்டி. "அட போ பாட்டி, வெயில் கொளுத்துது" என்று அவள் சொன்னதை அசட்டை செய்து விட்டு சென்றான் பேரன் ஹரி . இவர்களது பேச்சை கேட்டவாறே வீட்டினுள் சென்றாள் பேத்தி மீனா . ...

Read more

Bringing Python to Browser!

Jupyter notebook has always been a de-facto choice when you teach the Python programming language or while developing prototype machine learning models or doing exploratory data analysis. The reasons are manifold. The most important reason is due to its ability to interleave the rich set of explanatory notes using markdown cells and ...

Read more

Making Sense of Positional Encoding in Transformer

Motivation Are you wondering about the peculiar use of a sinusoidal function to encode the positional information in Transformer architecture? Are you asking why not just use simple one-hot encoding or something similar to encode positions?. Welcome, this article is for you. Perhaps you are here after reading a few articles explaining the posit...

Read more