08 - Convolution
Class: CSCE-421
Notes:
Pre:
- How to convert an image into a fix-length vector?
- We want this X to somehow remain the same if we rotate the input image
- Turns out the translation part is easy to do, the rotation part can be done but it is not easy to do
- Take an I image and turn it into a fix-length vector X
Image filtering
/CSCE-421/Ex2/Visual%20Aids/Pasted%20image%2020260219093618.png)
What is the idea of convolutional?
- An image is simply a two-dimensional data array with numbers
- If we have the 9-square box (3x3) as our filter (or sometimes called kernel) we will take this filter and do element-wise multiplication of some portion of the data array
- The output would be somehow a moving average of the inputs
- This translates to be a more blurry version of the image.
/CSCE-421/Ex2/Visual%20Aids/Pasted%20image%2020260219093639.png)
- Then we move one step forward (advance 1 position) so we can try to cover the whole image at some point
- Note, if we change out filter, our output can be very different
/CSCE-421/Ex2/Visual%20Aids/Pasted%20image%2020260219093700.png)
- We want this operation to somehow be related to the transformation of the image (what if we shift the location of the left box?)
- How the output would be changed in this case?
- The output will be just shifted also!
- This is called translation equivariant
- Means that if you do some transformation of the input, it will change the output equivalently
- How about rotation? if we rotate the object and apply the same filter, will the output be somehow related?
- In general the output will not be related to the original image at all
- What matters is the relative angle to the filter, if you rotate just the image multiplication is now messed up, so you are multiplicating together different numbers.
- If you rotate both the image and the filter you will then get the same output rotated.
- What if we apply 4 different kernels to the input image, each with a different angle of rotation. If the input image is rotated then you will just basically change the order of the output, but some of the angles might still be able to relate to the input image. Now you have a network that has rotation equivariance.
- But what happens if you rotate by just 3 degrees or an arbitrary number of degrees? So in general it is not that beneficial.
- Then it will break our system because we are only accounting for 4 different angles.
- This is not that intuitive to implement, to do this you have to rely in some transformations that might help do this.
- This is the idea of convolution
- Depending of a filter, the output will be very different
- The numbers on your filter are actually learnt from data, you will treat them as parameters and you will learn these parameters from data.
Box Filter
What does it do?
- Replaces each pixel with an average of its neighborhood
- Achieve smoothing effect (remove sharp features)
/CSCE-421/Ex2/Visual%20Aids/Pasted%20image%2020260219093857.png)
Smoothing with box filter
/CSCE-421/Ex2/Visual%20Aids/Pasted%20image%2020260219094504.png)
Practice with linear filters
/CSCE-421/Ex2/Visual%20Aids/Pasted%20image%2020260219094613.png)
- This filter would do nothing to the original image
/CSCE-421/Ex2/Visual%20Aids/Pasted%20image%2020260219094645.png)
/CSCE-421/Ex2/Visual%20Aids/Pasted%20image%2020260219094719.png)
- The right output is the Vertical Edge (absolute value)
- How does a negative number affect in this case?
- Remember we are doing element-wise multiplication
- This filter will cancel both sides of the vertical borders, not the horizontal border
/CSCE-421/Ex2/Visual%20Aids/Pasted%20image%2020260219095015.png)
- This is the Horizontal Edge (absolute value)
- If you make the filter slightly larger you can also do a 45 degreee detector and in principle any detector you can think of
- Somehow by just playing around with a filter we can modify our input image and make it easier to recognize it.
Image filtering
- Filters can be designed to detect edges of different orientations
- The detected edges can be combined to form object shapes, which are important for object recognition
- Convolutional neural networks are based on the idea of image convolutions
- Convolutional neural networks use data to train filter parameters
Summary
To understand Convolutional Neural Networks (CNNs), we first have to understand how computers see images and what a "convolution" actually is.
1. How Computers See Images When you look at a photograph, you see shapes and colors. A computer, however, simply sees a massive two-dimensional grid of numbers. Each number represents the brightness or color of a single pixel (e.g., 0 might be black, 255 might be white).
2. The Convolution Operation (Image Filtering) Imagine taking a tiny square grid, like a 3x3 box of numbers, and sliding it over your large image grid like a magnifying glass. This tiny box is called a filter or kernel.
Here is what happens at every step as you slide this filter across the image:
- You lay the 3x3 filter on top of a 3x3 patch of pixels in the top-left corner of the image.
- You do element-wise multiplication—you multiply the top-left number of the filter by the top-left pixel, the middle number by the middle pixel, and so on.
- Finally, you add all those 9 multiplied numbers together to get one single number.
- You record that number on a new, blank grid, shift your filter one step over, and repeat the process until you have covered the whole image.
The new grid you just created is the output. Depending on the exact numbers you put inside your 3x3 filter, this output image will look radically different.
3. Types of Filters Your notes show a few hand-crafted examples of what different numbers in a filter can do to an image:
- Box Filter (Smoothing): If you fill your 3x3 filter with equal fractions (like 1/9), the filter basically calculates the average brightness of the 9 pixels it is looking at. This replaces every pixel with the average of its neighborhood, which removes sharp details and creates a blurry, smoothing effect.
- Vertical Edge Detector: Imagine a filter where the left column is all -1s, the middle is 0s, and the right column is +1s. Because it multiplies the left pixels by negative numbers and adds them to the right pixels, it is basically calculating the difference between the left and right sides. If the image is a solid, flat color, the left and right sides cancel each other out to zero (black). But if there is a sharp transition from dark to light (a vertical edge), the math yields a huge number (white). This magically draws an outline over all the vertical edges in the image.
- Horizontal Edge Detector: By flipping those numbers so the top row is negative and the bottom row is positive, the filter cancels out vertical lines but perfectly highlights horizontal edges.
4. Equivariance (Translation vs. Rotation) This is a very important concept for your exam. We want our model to recognize an object no matter where it is.
- Translation Equivariance: "Translation" just means shifting an object left, right, up, or down. If a cat is in the top-left of the image, the filter will detect cat-edges in the top-left. If you shift the cat to the bottom-right, the filter detects the exact same edges, just shifted to the bottom-right. This is called Translation Equivariance: shifting the input shifts the output by the exact same equivalent amount.
- Rotation Equivariance: What if the cat is rotated upside down? Will the standard filters still work? No. If you rotate the image, the pixel math gets completely messed up relative to your filter (e.g., a vertical edge detector won't detect a horizontal edge). To fix this, you could apply 4 different filters rotated at different angles to catch the cat no matter its orientation. However, what if the cat is rotated by exactly 3 degrees? Accounting for every single arbitrary angle is computationally exhausting and breaks the system. Therefore, basic CNNs are naturally translation-equivariant, but not naturally rotation-equivariant.
5. The Magic of Deep Learning In the old days, computer scientists manually chose the 1s, 0s, and -1s in these filters to detect specific shapes. In Convolutional Neural Networks, we do not program the numbers in the filters. Instead, we treat the numbers inside the filters as learnable parameters (weights). The neural network uses training data to automatically learn and adjust the best numbers for these filters to extract whatever features (edges, textures, shapes) it needs to recognize objects.