Why images need special networks
A small 224 × 224 colour photo has 150,528 numbers. Connecting every pixel to every neuron of a normal neural network would need billions of weights — and it would have to re-learn “cat ear” separately for every position in the image.
Convolutional Neural Networks fix both problems with two ideas:
- Local patterns: look at small patches (e.g. 3×3 pixels) at a time.
- Weight sharing: use the same small filter at every position.
Convolution, step by step
A kernel (or filter) is a small grid of weights, for example a vertical-edge detector:
1 0 −1
1 0 −1
1 0 −1
Slide it over the image. At each position, multiply every pixel under the kernel by the weight on top of it and add everything up. That sum becomes one cell of the feature map.
- Where the image has a vertical edge (ink on the left, blank on the right), the sum is large and positive.
- On flat areas the positives and negatives cancel → 0.
In the 3D model, the yellow window slides across the image; each feature-map cell sticks out in proportion to its value (green positive, red negative).
Output size
For an n × n image and a k × k kernel with stride 1 and no padding, the output is (n − k + 1) × (n − k + 1). Our 8×8 image and 3×3 kernel give 6×6. Padding (adding a border of zeros) keeps the size the same; stride 2 skips every other position and halves it.
ReLU
Next, apply ReLU to every cell: negative values become 0. Only the “pattern found here” signals remain.
Pooling
Max-pooling takes each 2×2 block and keeps only its largest value, shrinking the map to a quarter of its size. This keeps the strongest signals, reduces computation, and makes the network less sensitive to small shifts of the object.
Stacking layers
A real CNN has many kernels per layer (each producing its own feature map) and many layers:
- early layers learn edges and colours,
- middle layers combine them into textures and parts (eyes, wheels),
- deep layers detect whole objects.
At the end, the maps are flattened into a vector and passed to fully connected layers that output class probabilities (via softmax). The kernels’ weights are learned with backpropagation — nobody hand-designs them.
Code
Convolution by hand with NumPy:
import numpy as np
image = np.array([[1,1,1,1,1,1,1,1],
[1,1,1,1,1,1,1,1],
[0,0,0,0,0,1,1,0],
[0,0,0,0,1,1,0,0],
[0,0,0,1,1,0,0,0],
[0,0,1,1,0,0,0,0],
[0,0,1,1,0,0,0,0],
[0,0,1,1,0,0,0,0]])
kernel = np.array([[1, 0, -1]] * 3) # vertical-edge detector
out = np.zeros((6, 6), dtype=int)
for r in range(6):
for c in range(6):
out[r, c] = np.sum(image[r:r+3, c:c+3] * kernel)
relu = np.maximum(out, 0)
pooled = relu.reshape(3, 2, 3, 2).max(axis=(1, 3)) # 2×2 max-pooling
print(out, pooled, sep="\n\n")
A tiny CNN in PyTorch for 28×28 digit images (MNIST):
import torch.nn as nn
model = nn.Sequential(
nn.Conv2d(1, 8, kernel_size=3), # 8 learned 3×3 kernels → 8 × 26 × 26
nn.ReLU(),
nn.MaxPool2d(2), # → 8 × 13 × 13
nn.Flatten(),
nn.Linear(8 * 13 * 13, 10), # 10 digit classes
)
Where are CNNs used?
Face unlock, medical scans (detecting tumours), self-driving cars, OCR / reading number plates, quality inspection in factories, satellite imagery — and as the vision part of many multimodal AI systems.
Common mistakes
- Getting the output size wrong — remember
n − k + 1(plus padding, divided by stride). - Thinking kernels are hand-made — in a trained CNN they are learned.
- Forgetting that each kernel spans all input channels (e.g. RGB = 3 channels).
Complexity at a glance
| Case / operation | Time | Why |
|---|---|---|
| One convolution (H×W image, k×k kernel) | O(H · W · k²) | |
| Parameters of a conv layer | k × k × channels_in × channels_out | Tiny compared with a fully connected layer. |
| Max-pooling 2×2 | O(H · W) | |
| Extra space | O(H · W) per feature map |
Quick check
Test yourself — pick an answer to see if you got it.
1. What does a kernel (filter) in a CNN do?
Each kernel is a small grid of learned weights; its response is high where the pattern appears.
2. An 8×8 image is convolved with a 3×3 kernel (stride 1, no padding). What is the size of the feature map?
Output size = 8 − 3 + 1 = 6 in each direction.
3. What does 2×2 max-pooling do to a 6×6 feature map?
Pooling shrinks the map and keeps the strongest responses, making the network more robust to small shifts.
4. Why do CNNs use far fewer weights than a fully connected network on images?
A 3×3 kernel has 9 weights no matter how big the image is.