Artificial Intelligence

Perceptron — The Artificial Neuron; Convolutional Neural Networks (CNN); Recurrent Neural Networks (RNN)

C-CAT

Perceptron — The Artificial Neuron

What is a Perceptron?

A perceptron is the simplest artificial neuron — the fundamental building block of all neural networks.

Invented by Frank Rosenblatt in 1958.

Perceptron Model

     x1 ---w1--->+
                 |
     x2 ---w2--->+------> Sigma(wi*xi + b) ----> Activation f ----> Output (y_hat)
                 |
     x3 ---w3--->+
          ^
         bias (b)

Mathematical Formula

Net input:  z = w1*x1 + w2*x2 + w3*x3 + b  =  Sigma(wi*xi) + b

Step function output:
  y = f(z) = 1   if z >= 0
            = 0   if z < 0

Where:

  • x = input values (features)
  • w (Weights) = determine the influence of each input; adjusted during training
  • b (Bias) = allows the perceptron to adjust its threshold independently of inputs; improves flexibility

Single-Layer Perceptron

  • Can solve linearly separable problems only
  • Can implement AND and OR logic gates
  • Cannot solve XOR problem (not linearly separable)

AND Gate Truth Table with Perceptron:

x1  x2  | Target | Sum (w1=0.5, w2=0.5, b=-0.8) | Output
 0   0  |   0    | 0 - 0.8 = -0.8               |   0   ok
 0   1  |   0    | 0.5 - 0.8 = -0.3             |   0   ok
 1   0  |   0    | 0.5 - 0.8 = -0.3             |   0   ok
 1   1  |   1    | 1.0 - 0.8 = +0.2             |   1   ok

Multi-Layer Perceptron (MLP)

  • Has multiple layers: input → hidden → output
  • Can solve non-linear problems including XOR
  • Trained using backpropagation algorithm
  • Adding hidden layers = "going deeper" = Deep Learning

XOR Problem — Why MLP is needed:

x1  x2  | XOR Output | NOT linearly separable
 0   0  |     0      | Cannot draw a single straight line to separate 0s from 1s
 0   1  |     1      | → Need a curved boundary → MLP with hidden layers
 1   0  |     1      |
 1   1  |     0      |

Convolutional Neural Networks (CNN)

Why CNN?

Fully-connected networks don't work well for images:

  • A 224x224 RGB image has 224 x 224 x 3 = 150,528 inputs
  • Connecting to 1000 neurons would need 150 million parameters → overfit

CNNs use local connections and parameter sharing to dramatically reduce parameters.

CNN Architecture

Input Image
    |
Conv Layer (Feature maps)
    |
Activation (ReLU)
    |
Pooling Layer (Downsample)
    |
[Repeat Conv-ReLU-Pool blocks...]
    |
Flatten (to 1D vector)
    |
Fully Connected Layers
    |
Softmax Output

CNN Layer Types

LayerFunction
Convolutional LayerApplies filters to detect features (edges, textures, shapes)
ReLU ActivationIntroduces non-linearity, removes negatives
Pooling Layer (Max/Avg)Reduces spatial dimensions; provides translation invariance
FlattenConverts 2D feature maps to 1D vector
Fully Connected (FC)Standard layers for final classification

Applications of CNN

  • Image Classification — what is in this image?
  • Object Detection — where are objects in this image? (YOLO, Faster R-CNN)
  • Face Recognition — identifies persons (FaceNet, DeepFace by Facebook)
  • Medical Imaging — tumor detection in X-ray/MRI scans

Autonomous Vehicles — lane detection, pedestrian recognition

  • OCR — reading text from images

Recurrent Neural Networks (RNN)

Why RNN?

Standard neural networks treat each input independently. But sequential data (text, speech, time series) has order and context that must be captured.

Examples:

  • "The cat sat on the mat" — word order matters
  • Stock prices over time — past affects future
  • Speech signal — phonemes have temporal dependencies

RNNs preserve context through a hidden state that acts as memory.

RNN Architecture

Input: "The cat sat on the mat"

x1="The" --> [h1] --> y1
              |
x2="cat" --> [h2] --> y2
              |
x3="sat" --> [h3] --> y3
              ...

The hidden state h carries context from previous words.

Problem: Vanishing Gradient

Basic RNNs suffer from vanishing gradients — gradients become extremely small during backpropagation, making it hard to learn long-term dependencies.

Solution: LSTM (Long Short-Term Memory)

LSTM Gates

GateFunction
Forget GateDecides what information to discard from cell state
Input GateDecides what new information to store in cell state
Output GateDecides what part of cell state to output

Applications of RNN/LSTM

  • Machine Translation — "Bonjour" → "Hello"
  • Speech Recognition — audio waveform → text
  • Text Generation — autocomplete, story generation
  • Sentiment Analysis — positive/negative classification
  • Time Series — stock prices, weather, energy consumption
  • Video Analysis — action recognition in video clips

Continue learning

Related notes

Put this topic into timed practice

Open mock tests when you want full-exam pacing, or keep drilling in practice mode.