Artificial Intelligence
Machine Learning — Unsupervised Learning; Machine Learning — Reinforcement Learning
C-CAT
Machine Learning — Unsupervised Learning
What is Unsupervised Learning?
Unsupervised Learning is where you have only input data (X) — no corresponding output labels (Y).
The goal is to model the underlying structure or distribution in the data to learn more about the data.
Why "Unsupervised"? There are no correct answers and no teacher. Algorithms are left to discover interesting structure in the data on their own.
Structure discovered in the form of:
- Groups / Clusters — natural groupings in data
- Associations — items that frequently co-occur
Mostly used for EDA (Exploratory Data Analysis)
Process:
Unlabeled Data (X only)
↓
Algorithm discovers patterns
↓
Groups / Clusters / Associations
↓
Business insights
Characteristics
- Analyzes and clusters unlabeled datasets
- Finds hidden patterns and data without human intervention
- Training model has only input parameter values
- Discovers groups or patterns on its own
11.1 Clustering Problems
Clustering discovers the inherent groupings in data, such as grouping customers by purchasing behavior.
- Batsman vs Bowler — classify cricketers based on performance stats
- Customer segmentation — customers spending more vs less money
Additional Examples:
- Document grouping — cluster news articles by topic
- Image segmentation — group pixels into regions
- Gene expression analysis — group genes with similar expression
Clustering Algorithms:
| Algorithm | How it Works | Pros | Cons |
|---|---|---|---|
| K-Means | Assigns points to k centroids; iteratively updates | Fast, simple | Must specify k |
| Hierarchical Clustering | Builds dendrogram of clusters | No need to specify k | Slow for large data |
| DBSCAN | Density-based; finds arbitrary shapes | No k needed; handles noise | Sensitive to parameters |
K-Means Algorithm Steps:
1. Choose k (number of clusters)
2. Initialize k centroids randomly
3. Assign each point to nearest centroid
4. Recalculate centroids as mean of assigned points
5. Repeat steps 3-4 until convergence (no changes)
11.2 Association Problems
Association discovers rules describing large portions of data — items that frequently appear together.
Classic Example — Market Basket Analysis:
- People who buy bread also tend to buy butter
- People who buy diapers also tend to buy beer
Association Rule Format:
{Antecedent} --> {Consequent}
{bread} --> {butter} [support: 60%, confidence: 80%, lift: 1.5]
Metrics:
- Support — How frequently the itemset appears in transactions
Confidence — How often the rule is correct (P(butter | bread))
- Lift — How much better than random chance the rule is
Algorithms:
| Algorithm | Description |
|---|---|
| Apriori | Generates frequent itemsets bottom-up; prunes infrequent sets |
| Eclat | Uses vertical data format; faster than Apriori |
| FP-Growth | Uses tree structure; avoids candidate generation |
Applications:
- Supermarket product placement
- E-commerce recommendation systems
Medical symptom correlation analysis
- Web page clickstream analysis
Machine Learning — Reinforcement Learning
What is Reinforcement Learning?
Reinforcement Learning (RL) is about taking suitable actions to maximize reward in a particular situation.
It is employed by software and machines to find the best possible behavior or path to take in a specific situation.
Key Difference from Supervised Learning
In Supervised Learning: Training data has the answer key — model is trained with correct answers.
In Reinforcement Learning: There is no answer key — the reinforcement agent decides what to do to perform the given task. In the absence of a training dataset, it learns from its own experience.
RL Framework
+--------------------------------------------------+
| ENVIRONMENT |
| |
| State (s) ---------> AGENT |
| | |
| Reward (r) <------ Action (a) |
| |
+--------------------------------------------------+
Key Terms:
- Agent — the learner / decision-maker
- Environment — what the agent interacts with
- State (s) — current situation of the agent
- Action (a) — what the agent does in a state
- Reward (r) — positive or negative feedback signal
- Policy (pi) — agent's strategy for choosing actions
- Value Function — expected future cumulative reward
Examples
- Resource management in computer clusters — optimize CPU/memory allocation
Traffic Light Control — optimize signal timings to reduce congestion 3. Robotics — teach robots to walk, grasp, balance 4. Web system configuration — auto-tune system parameters 5. Chemistry — optimize reaction conditions
RL Algorithms
| Algorithm | Description | Use Case |
|---|---|---|
| Q-Learning | Model-free; learns action-value (Q) function | Game playing, discrete actions |
| Deep Q-Learning (DQN) | Neural network approximates Q-function | Atari games, complex state spaces |
| Policy Gradient | Directly optimizes policy | Continuous action spaces |
| Actor-Critic | Combines policy gradient and value function | Complex environments |
| PPO (Proximal Policy Optimization) | Stable, efficient policy gradient | Most modern RL applications |
Famous RL Achievements
| System | Task | Result |
|---|---|---|
| DeepMind AlphaGo | Board game Go | Defeated world champion Lee Sedol (2016) |
| OpenAI Five | Dota 2 (MOBA game) | Defeated professional human teams |
| AlphaStar | StarCraft II | Reached Grandmaster level |
| AlphaFold 2 | Protein structure | Solved 50-year biology problem |
Continue learning
Related notes
Put this topic into timed practice
Open mock tests when you want full-exam pacing, or keep drilling in practice mode.