Model Training Strategies
Training large language models from scratch is the job of tech giants. Often, the pre-trained proprietary models are adapted to downstream tasks using instruction fine-tuning. However, doing full fine-tuning of model parameters increases the model performance. Of course, full fine- tuning of large models with Billions of parameters requires a go...
Introduction to Large Language Models - Course
For the past three months, I have been quite busy building course materials (lecture slides, graded assignments and coding assignments) for the first offering of the course Introduction to Large Language Models by Prof.Mitesh Khapra. It has been challenging work as we have committed to offer the course in the JAN 2024 term. Every challenge is an...
Positional Encoding In Transformers
Introduction
One question that bothers our mind when we read positional encoding is whether it helps the model or not. It is observed that adding position embeddings only helps marginally in a convolutional neural network as CNN uses relative position embedding (implicitly). It is not the case for transformers as the model is permutation invari...
Data Pipeline for Large Language models
Data is the fuel for any machine learning model despite the type of learning algorithm (Gradient-based or tree-based) being used. To train and test the model’s generalization capacity, typically, we divide the available samples into three sets: Training, Validation and Test. The typical requirement is that the samples in the test set should be d...
Experimental Settings of Famous Language Models
GPT (Generative Pre-trained Transformer)
Pre-training Dataset: Book corpus (0.8 Billion words)
Unsupervised objective: CLM (Autoregressive)
Tokenizer: Byte Pair Encoding (BPE)
Vocab size: 40K
Architecture: Decoder only (12 Layers)
Activation: GELU
Attention: Dense
FFN: Dense
Attention mask: Causal Mask
Positional ...
Emergence of Large Language Models (LLMs)
Motivation
Usually, in traditional machine learning, we use numerous approaches (model selection) like $K-$fold cross-validation and grid search to find the best model that generalizes well in the real world. However, when it comes to deep learning, it is quite challenging due to compute-cost constraints. It holds for neural language models too....
Lagrange Multiplier : Intuition via Interaction
Introduction
I guess you end up being here after coming across the term “constrained optimization” or “Lagrangian” and wanted to understand what “Lagrange multiplier is?”. Well, in this post, I help you understand the foundation of it with interactive plots (you can find plenty of mathematical reasoning on the net). Let’s get straight to the poi...
Maximum Likelihood Estimation
Introduction
The concept of estimation of an unknown quantity from the given observations has been a fascinating area of study for many centuries. However, it remains elusive for many beginners. Let’s start with a concept that we are already familiar with. Here is a sequence $x_1=[1,3,5,7,9,11, \times,\cdots,]$. What could be the value of the se...
33 post articles, 5 pages.