Artificial Intelligence
Perceptron — The Artificial Neuron; Convolutional Neural Networks (CNN); Recurrent Neural Networks (RNN)
C-CAT
Perceptron — The Artificial Neuron
What is a Perceptron?
A perceptron is the simplest artificial neuron — the fundamental building block of all neural networks.
Invented by Frank Rosenblatt in 1958.
Perceptron Model
x1 ---w1--->+
|
x2 ---w2--->+------> Sigma(wi*xi + b) ----> Activation f ----> Output (y_hat)
|
x3 ---w3--->+
^
bias (b)
Mathematical Formula
Net input: z = w1*x1 + w2*x2 + w3*x3 + b = Sigma(wi*xi) + b
Step function output:
y = f(z) = 1 if z >= 0
= 0 if z < 0
Where:
- x = input values (features)
- w (Weights) = determine the influence of each input; adjusted during training
- b (Bias) = allows the perceptron to adjust its threshold independently of inputs; improves flexibility
Single-Layer Perceptron
- Can solve linearly separable problems only
- Can implement AND and OR logic gates
- Cannot solve XOR problem (not linearly separable)
AND Gate Truth Table with Perceptron:
x1 x2 | Target | Sum (w1=0.5, w2=0.5, b=-0.8) | Output
0 0 | 0 | 0 - 0.8 = -0.8 | 0 ok
0 1 | 0 | 0.5 - 0.8 = -0.3 | 0 ok
1 0 | 0 | 0.5 - 0.8 = -0.3 | 0 ok
1 1 | 1 | 1.0 - 0.8 = +0.2 | 1 ok
Multi-Layer Perceptron (MLP)
- Has multiple layers: input → hidden → output
- Can solve non-linear problems including XOR
- Trained using backpropagation algorithm
- Adding hidden layers = "going deeper" = Deep Learning
XOR Problem — Why MLP is needed:
x1 x2 | XOR Output | NOT linearly separable
0 0 | 0 | Cannot draw a single straight line to separate 0s from 1s
0 1 | 1 | → Need a curved boundary → MLP with hidden layers
1 0 | 1 |
1 1 | 0 |
Convolutional Neural Networks (CNN)
Why CNN?
Fully-connected networks don't work well for images:
- A 224x224 RGB image has 224 x 224 x 3 = 150,528 inputs
- Connecting to 1000 neurons would need 150 million parameters → overfit
CNNs use local connections and parameter sharing to dramatically reduce parameters.
CNN Architecture
Input Image
|
Conv Layer (Feature maps)
|
Activation (ReLU)
|
Pooling Layer (Downsample)
|
[Repeat Conv-ReLU-Pool blocks...]
|
Flatten (to 1D vector)
|
Fully Connected Layers
|
Softmax Output
CNN Layer Types
| Layer | Function |
|---|---|
| Convolutional Layer | Applies filters to detect features (edges, textures, shapes) |
| ReLU Activation | Introduces non-linearity, removes negatives |
| Pooling Layer (Max/Avg) | Reduces spatial dimensions; provides translation invariance |
| Flatten | Converts 2D feature maps to 1D vector |
| Fully Connected (FC) | Standard layers for final classification |
Applications of CNN
- Image Classification — what is in this image?
- Object Detection — where are objects in this image? (YOLO, Faster R-CNN)
- Face Recognition — identifies persons (FaceNet, DeepFace by Facebook)
- Medical Imaging — tumor detection in X-ray/MRI scans
Autonomous Vehicles — lane detection, pedestrian recognition
- OCR — reading text from images
Recurrent Neural Networks (RNN)
Why RNN?
Standard neural networks treat each input independently. But sequential data (text, speech, time series) has order and context that must be captured.
Examples:
- "The cat sat on the mat" — word order matters
- Stock prices over time — past affects future
- Speech signal — phonemes have temporal dependencies
RNNs preserve context through a hidden state that acts as memory.
RNN Architecture
Input: "The cat sat on the mat"
x1="The" --> [h1] --> y1
|
x2="cat" --> [h2] --> y2
|
x3="sat" --> [h3] --> y3
...
The hidden state h carries context from previous words.
Problem: Vanishing Gradient
Basic RNNs suffer from vanishing gradients — gradients become extremely small during backpropagation, making it hard to learn long-term dependencies.
Solution: LSTM (Long Short-Term Memory)
LSTM Gates
| Gate | Function |
|---|---|
| Forget Gate | Decides what information to discard from cell state |
| Input Gate | Decides what new information to store in cell state |
| Output Gate | Decides what part of cell state to output |
Applications of RNN/LSTM
- Machine Translation — "Bonjour" → "Hello"
- Speech Recognition — audio waveform → text
- Text Generation — autocomplete, story generation
- Sentiment Analysis — positive/negative classification
- Time Series — stock prices, weather, energy consumption
- Video Analysis — action recognition in video clips
Continue learning
Related notes
Put this topic into timed practice
Open mock tests when you want full-exam pacing, or keep drilling in practice mode.