Artificial Intelligence

Deep Learning; Neural Networks — Architecture & Types

C-CAT

Deep Learning

What is Deep Learning?

Deep Learning is a method in AI that teaches computers to process data in a way that is inspired by the human brain.

Deep learning models can recognize complex patterns in pictures, text, sounds and other data to produce accurate insights and predictions.

AI attempts to train computers to think and learn as humans do. Deep learning is the most powerful current tool to achieve this.

Deep Learning Applications

  • Digital assistants — Alexa, Siri, Google Assistant
  • Voice-activated television remotes — speech recognition
  • Fraud detection — banking and financial systems

Automatic facial recognition — phone unlock, security cameras

Why "Deep"?

"Deep" refers to the multiple hidden layers in a neural network:

  • Shallow network: 1-2 layers
  • Deep network: 3+ layers (often dozens or hundreds)

What each layer learns:

(For image recognition:)
Layer 1: Edges and gradients
Layer 2: Simple shapes (corners, curves)
Layer 3: Complex shapes (eyes, ears, wheels)
Layer 4: Complete objects (faces, cars)

How Deep Learning Works

Deep learning algorithms are neural networks modeled after the human brain:

  • Human brain contains millions of interconnected neurons working together to learn and process information.
  • Similarly, deep learning neural networks are made of many layers of artificial neurons working together.
  • Each neuron in one layer connects to neurons in the next layer.

Deep Learning Training Process

Step 1: Forward Pass
  Input → Hidden Layers → Output → Prediction

Step 2: Calculate Loss
  Compare prediction with correct answer
  Loss = error measure
(e.g., MSE for regression, Cross-entropy for classification)

Step 3: Backward Pass (Backpropagation)
  Compute gradients of loss with respect to each
weight
  Chain rule: dL/dw = dL/dy x dy/dw

Step 4: Update Weights (Gradient Descent)
  w = w - alpha x (dL/dw)
  alpha = learning rate

Repeat until convergence (loss stops decreasing)

Key Hyperparameters

HyperparameterDescriptionTypical Values
Learning Rate (alpha)Controls weight update step size0.001 – 0.1
Batch SizeNumber of samples per weight update32, 64, 128
EpochsNumber of complete passes through training data10 – 1000
LayersNumber of hidden layers3 – 100+
Neurons per LayerWidth of each layer64 – 4096
Dropout RateFraction of neurons randomly disabled0.2 – 0.5

Neural Networks — Architecture & Types

Biological Neuron — The Inspiration

A neuron (nerve cell) is an electrically excitable cell that communicates with other cells via specialized connections called synapses.

Components:

                 Dendrites
                 (receive signals)
                     |
                     v
              +-------------+
              |  Soma       |
              | (Cell Body) |----> Axon ----> Axon Terminal
              | Processes   |                (sends signals)
              | signals     |
              +-------------+
  • Soma (Cell Body) — processes incoming information
  • Dendrites — cellular extensions that receive signals from other neurons
  • Axon — carries signals away from the cell body to other neurons
  • Synapses — junctions between neurons; where signal transfer occurs
  • Neurotransmitters — chemicals carrying signals across synapses

Artificial Neural Network (ANN)

An ANN mimics the biological neural network with:

  • Artificial neurons (nodes)

Weighted connections (edges)

  • Activation functions (mimic neuron firing thresholds)

Layers of a Neural Network

INPUT LAYER        HIDDEN LAYERS              OUTPUT LAYER
   (x1)               (h1) (h2)                 (y)
   (x2)  ------------> (h3) (h4) ------------> (Result)
   (x3)               (h5) (h6)
LayerRole
Input LayerReceives raw input features; passes to first hidden layer
Hidden Layer(s)Performs weighted sums and activation; learns abstract features
Output LayerProduces final prediction (class probability or value)

"Deep" = Multiple Hidden Layers (3+)

Characteristics of ANN

  • Neurally-implemented mathematical model
  • Contains huge number of interconnected processing elements (neurons)
  • Information stored as weighted linkages between neurons
  • Input signals arrive through connections with associated weights
  • Can learn, recall and generalize
  • No single neuron carries specific information — collective behavior determines output
  • Modifies connection weights through learning (training)

Types of Neural Networks

TypeArchitectureBest ForExample
ANN (MLP)Fully connected layersGeneral classification/regressionTabular data
CNNConvolutional + pooling layersImages, videosResNet, VGG
RNNRecurrent loopsSequences, text, time seriesLanguage models
LSTMRNN with memory gatesLong sequencesTranslation
GANGenerator + DiscriminatorImage generationDALL-E
TransformerSelf-attention mechanismNLP, modern AIBERT, GPT
AutoencoderEncoder + DecoderCompression, anomaly detectionVariational AE

Activation Functions

Activation functions introduce non-linearity into neural networks, enabling them to learn complex patterns:

FunctionFormulaRangeBest Use
Sigmoid1/(1+e^-x)(0, 1)Binary output
Tanh(e^x - e^-x)/(e^x + e^-x)(-1, 1)Hidden layers
ReLUmax(0, x)[0, inf)Most hidden layers
Leaky ReLUmax(0.01x, x)(-inf, inf)Fixes "dying ReLU"
Softmaxe^xi / sum(e^xj)(0,1); sums to 1Multi-class output

Continue learning

Related notes

Put this topic into timed practice

Open mock tests when you want full-exam pacing, or keep drilling in practice mode.