Home

Model Training Strategies

Training large language models from scratch is the job of tech giants. Often, the pre-trained proprietary models are adapted to downstream tasks using instruction fine-tuning. However, doing full fine-tuning of model parameters increases the model performance. Of course, full fine- tuning of large models with Billions of parameters requires a go...

Read more

Introduction to Large Language Models - Course

For the past three months, I have been quite busy building course materials (lecture slides, graded assignments and coding assignments) for the first offering of the course Introduction to Large Language Models by Prof.Mitesh Khapra. It has been challenging work as we have committed to offer the course in the JAN 2024 term. Every challenge is an...

Read more

Positional Encoding In Transformers

Introduction One question that bothers our mind when we read positional encoding is whether it helps the model or not. It is observed that adding position embeddings only helps marginally in a convolutional neural network as CNN uses relative position embedding (implicitly). It is not the case for transformers as the model is permutation invari...

Read more

Data Pipeline for Large Language models

Data is the fuel for any machine learning model despite the type of learning algorithm (Gradient-based or tree-based) being used. To train and test the model’s generalization capacity, typically, we divide the available samples into three sets: Training, Validation and Test. The typical requirement is that the samples in the test set should be d...

Read more

Experimental Settings of Famous Language Models

GPT (Generative Pre-trained Transformer) Pre-training Dataset: Book corpus (0.8 Billion words) Unsupervised objective: CLM (Autoregressive) Tokenizer: Byte Pair Encoding (BPE) Vocab size: 40K Architecture: Decoder only (12 Layers) Activation: GELU Attention: Dense FFN: Dense Attention mask: Causal Mask Positional ...

Read more

Emergence of Large Language Models (LLMs)

Motivation Usually, in traditional machine learning, we use numerous approaches (model selection) like $K-$fold cross-validation and grid search to find the best model that generalizes well in the real world. However, when it comes to deep learning, it is quite challenging due to compute-cost constraints. It holds for neural language models too....

Read more

Lagrange Multiplier : Intuition via Interaction

Introduction I guess you end up being here after coming across the term “constrained optimization” or “Lagrangian” and wanted to understand what “Lagrange multiplier is?”. Well, in this post, I help you understand the foundation of it with interactive plots (you can find plenty of mathematical reasoning on the net). Let’s get straight to the poi...

Read more

Maximum Likelihood Estimation

Introduction The concept of estimation of an unknown quantity from the given observations has been a fascinating area of study for many centuries. However, it remains elusive for many beginners. Let’s start with a concept that we are already familiar with. Here is a sequence $x_1=[1,3,5,7,9,11, \times,\cdots,]$. What could be the value of the se...

Read more