Artificial Intelligence
Deep Learning; Neural Networks — Architecture & Types
C-CAT
Deep Learning
What is Deep Learning?
Deep Learning is a method in AI that teaches computers to process data in a way that is inspired by the human brain.
Deep learning models can recognize complex patterns in pictures, text, sounds and other data to produce accurate insights and predictions.
AI attempts to train computers to think and learn as humans do. Deep learning is the most powerful current tool to achieve this.
Deep Learning Applications
- Digital assistants — Alexa, Siri, Google Assistant
- Voice-activated television remotes — speech recognition
- Fraud detection — banking and financial systems
Automatic facial recognition — phone unlock, security cameras
Why "Deep"?
"Deep" refers to the multiple hidden layers in a neural network:
- Shallow network: 1-2 layers
- Deep network: 3+ layers (often dozens or hundreds)
What each layer learns:
(For image recognition:)
Layer 1: Edges and gradients
Layer 2: Simple shapes (corners, curves)
Layer 3: Complex shapes (eyes, ears, wheels)
Layer 4: Complete objects (faces, cars)
How Deep Learning Works
Deep learning algorithms are neural networks modeled after the human brain:
- Human brain contains millions of interconnected neurons working together to learn and process information.
- Similarly, deep learning neural networks are made of many layers of artificial neurons working together.
- Each neuron in one layer connects to neurons in the next layer.
Deep Learning Training Process
Step 1: Forward Pass
Input → Hidden Layers → Output → Prediction
Step 2: Calculate Loss
Compare prediction with correct answer
Loss = error measure
(e.g., MSE for regression, Cross-entropy for classification)
Step 3: Backward Pass (Backpropagation)
Compute gradients of loss with respect to each
weight
Chain rule: dL/dw = dL/dy x dy/dw
Step 4: Update Weights (Gradient Descent)
w = w - alpha x (dL/dw)
alpha = learning rate
Repeat until convergence (loss stops decreasing)
Key Hyperparameters
| Hyperparameter | Description | Typical Values |
|---|---|---|
| Learning Rate (alpha) | Controls weight update step size | 0.001 – 0.1 |
| Batch Size | Number of samples per weight update | 32, 64, 128 |
| Epochs | Number of complete passes through training data | 10 – 1000 |
| Layers | Number of hidden layers | 3 – 100+ |
| Neurons per Layer | Width of each layer | 64 – 4096 |
| Dropout Rate | Fraction of neurons randomly disabled | 0.2 – 0.5 |
Neural Networks — Architecture & Types
Biological Neuron — The Inspiration
A neuron (nerve cell) is an electrically excitable cell that communicates with other cells via specialized connections called synapses.
Components:
Dendrites
(receive signals)
|
v
+-------------+
| Soma |
| (Cell Body) |----> Axon ----> Axon Terminal
| Processes | (sends signals)
| signals |
+-------------+
- Soma (Cell Body) — processes incoming information
- Dendrites — cellular extensions that receive signals from other neurons
- Axon — carries signals away from the cell body to other neurons
- Synapses — junctions between neurons; where signal transfer occurs
- Neurotransmitters — chemicals carrying signals across synapses
Artificial Neural Network (ANN)
An ANN mimics the biological neural network with:
- Artificial neurons (nodes)
Weighted connections (edges)
- Activation functions (mimic neuron firing thresholds)
Layers of a Neural Network
INPUT LAYER HIDDEN LAYERS OUTPUT LAYER
(x1) (h1) (h2) (y)
(x2) ------------> (h3) (h4) ------------> (Result)
(x3) (h5) (h6)
| Layer | Role |
|---|---|
| Input Layer | Receives raw input features; passes to first hidden layer |
| Hidden Layer(s) | Performs weighted sums and activation; learns abstract features |
| Output Layer | Produces final prediction (class probability or value) |
"Deep" = Multiple Hidden Layers (3+)
Characteristics of ANN
- Neurally-implemented mathematical model
- Contains huge number of interconnected processing elements (neurons)
- Information stored as weighted linkages between neurons
- Input signals arrive through connections with associated weights
- Can learn, recall and generalize
- No single neuron carries specific information — collective behavior determines output
- Modifies connection weights through learning (training)
Types of Neural Networks
| Type | Architecture | Best For | Example |
|---|---|---|---|
| ANN (MLP) | Fully connected layers | General classification/regression | Tabular data |
| CNN | Convolutional + pooling layers | Images, videos | ResNet, VGG |
| RNN | Recurrent loops | Sequences, text, time series | Language models |
| LSTM | RNN with memory gates | Long sequences | Translation |
| GAN | Generator + Discriminator | Image generation | DALL-E |
| Transformer | Self-attention mechanism | NLP, modern AI | BERT, GPT |
| Autoencoder | Encoder + Decoder | Compression, anomaly detection | Variational AE |
Activation Functions
Activation functions introduce non-linearity into neural networks, enabling them to learn complex patterns:
| Function | Formula | Range | Best Use |
|---|---|---|---|
| Sigmoid | 1/(1+e^-x) | (0, 1) | Binary output |
| Tanh | (e^x - e^-x)/(e^x + e^-x) | (-1, 1) | Hidden layers |
| ReLU | max(0, x) | [0, inf) | Most hidden layers |
| Leaky ReLU | max(0.01x, x) | (-inf, inf) | Fixes "dying ReLU" |
| Softmax | e^xi / sum(e^xj) | (0,1); sums to 1 | Multi-class output |
Continue learning
Related notes
Put this topic into timed practice
Open mock tests when you want full-exam pacing, or keep drilling in practice mode.