08 - Convolution

Class: CSCE-421


Notes:

Pre:

Image filtering

Pasted image 20260219093618.png500

What is the idea of convolutional?

Pasted image 20260219093639.png500

Pasted image 20260219093700.png500

Box Filter

What does it do?

Pasted image 20260219093857.png150

Smoothing with box filter

Pasted image 20260219094504.png400

Practice with linear filters

Pasted image 20260219094613.png500

Pasted image 20260219094645.png500

Pasted image 20260219094719.png500

Pasted image 20260219095015.png500

Image filtering

Summary

To understand Convolutional Neural Networks (CNNs), we first have to understand how computers see images and what a "convolution" actually is.

1. How Computers See Images When you look at a photograph, you see shapes and colors. A computer, however, simply sees a massive two-dimensional grid of numbers. Each number represents the brightness or color of a single pixel (e.g., 0 might be black, 255 might be white).

2. The Convolution Operation (Image Filtering) Imagine taking a tiny square grid, like a 3x3 box of numbers, and sliding it over your large image grid like a magnifying glass. This tiny box is called a filter or kernel.

Here is what happens at every step as you slide this filter across the image:

The new grid you just created is the output. Depending on the exact numbers you put inside your 3x3 filter, this output image will look radically different.

3. Types of Filters Your notes show a few hand-crafted examples of what different numbers in a filter can do to an image:

4. Equivariance (Translation vs. Rotation) This is a very important concept for your exam. We want our model to recognize an object no matter where it is.

5. The Magic of Deep Learning In the old days, computer scientists manually chose the 1s, 0s, and -1s in these filters to detect specific shapes. In Convolutional Neural Networks, we do not program the numbers in the filters. Instead, we treat the numbers inside the filters as learnable parameters (weights). The neural network uses training data to automatically learn and adjust the best numbers for these filters to extract whatever features (edges, textures, shapes) it needs to recognize objects.