Close Menu
AI News TodayAI News Today

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    Qwen3.8 27B addition in words

    Connecting AI agents to enterprise knowledge

    OpenAI launches visual ads that appear alongside image generation results

    Facebook X (Twitter) Instagram
    • About Us
    • Contact Us
    Facebook X (Twitter) Instagram Pinterest Vimeo
    AI News TodayAI News Today
    • Home
    • AI News
    • AI Reviews
    • AI Tools
    • AI Tutorials
    • Chatbots
    • Free AI Tools
    • Artificial Intelligence
    AI News TodayAI News Today
    Home»AI Tools»Computer Vision: SIFT algorithm (Scale Invariant Feature Transform)
    AI Tools

    Computer Vision: SIFT algorithm (Scale Invariant Feature Transform)

    By No Comments13 Mins Read
    Share Facebook Twitter Pinterest LinkedIn Tumblr Reddit Telegram Email
    Computer Vision: SIFT algorithm (Scale Invariant Feature Transform)
    Share
    Facebook Twitter LinkedIn Pinterest Email

    SIFT is one of the most widely known algorithms in computer vision. Its core objective consists of detecting object keypoints, generating descriptors for them, and matching the same objects across images.

    As the name suggests, SIFT is a scale-invariant algorithm, meaning that the same object can appear at different scales in a pair of images, and SIFT will still be able to successfully detect its keypoints.

    In addition, SIFT is rotation-invariant, making matching possible for rotated objects as well.

    Let us take a closer look at how SIFT works under the hood.

    Note: In this article, we will refer to the Laplacian of Gaussian (LoG) as a transformation used for edge detection in images. If you are unfamiliar with this technique, it is recommended that you go through one of the edge detection articles.

    In its workflow, SIFT constructs several versions of the original image by applying resize and Gaussian blur transformations.

    For simplicity, let’s imagine that I(x, y) is an original image. First, with chosen values of k and σ1, SIFT constructs several versions of the original image by applying Gaussian smoothing with different standard deviations: σ1, k⋅σ1, k2⋅σ1, k3⋅σ1, … , where k > 1.

    This results in a sequence of images in which each subsequent image is slightly blurrier than the previous one. This sequence of images is called an octave.

    Then SIFT computes the pairwise differences D1, D2, …, Dn between the resulting images, known as the difference of Gaussians (DoG). These differences highlight pixels with high intensity changes. After that, the algorithm stacks the Di and tries to find local extrema in them. Here is how it is done:

    For each point in Di(x, y), SIFT examines its 26 neighbours:

    • 8 adjacent points on Di level;

    • 9 points directly above Di(x, y) (on the Di+1 level);

    • 9 points directly below Di(x, y) (on the Di-1 level);

    Then one of the following three cases is possible:

    • If Di(x, y) is greater than all of its 26 neighbouring points, then SIFT marks it as a maximum.

    • If Di(x, y) is less than all 26 neighboring points, then SIFT marks it as a minimum;

    • Otherwise, the point Di(x, y) is skipped.

    For simplicity, the point Di(x, y) and its 26 neighboring points can be visualized as a 3x3x3 grid with the center at Di(x, y). This procedure allows the identification of the strongest features.

    The found extrema values represent points of interest. In fact, there can be too many of them; that is why SIFT applies thresholding or another operator to retain only those that represent the greatest changes.

    To account for different scale variations, the same process is repeated for an initial image reduced (downsampled) in width and height by a factor of two. As a result, a new octave sequence is constructed with greater Gaussian noise applied to it, having the following σ values: σ2, k⋅σ2, k2⋅σ1, k3⋅σ2, … , where σ2 = 2σ1. As before, extrema values are found from image differences using the 3x3x3 grid method.

    An example of two constructed octaves. Each octave contains 5 images with progressively increasing blur levels. The first (the lowest) image in the second octave is a logical continuation of the last (the highest) image in the first octave. While it might have a lower value of σ, this is compensated for by the image’s downsampled size.

    For the third iteration, the image is downsampled again (reduced in width and height by a factor of two), and a new, blurrier octave is constructed with σ values as σ3, k⋅σ3, k2⋅σ3, k3⋅σ3, … , where σ3 = 2σ2 = 4σ1.

    The entire process is repeated for a specified number of iterations.

    We understood how to find points of interest. Let’s now answer several important questions to build intuition about the process.

    Why DoG instead of LoG?

    In the past, we found that the Laplacian of Gaussian (LoG) is a very useful transformation for identifying edges in images. At the same time, it turns out that there exists a very good approximation for the difference of two scaled LoGs applied to the same image:

    DoG = nkσ – nσ ≈ (k – 1)σ2 ⋅ ▽2nσ

    In fact, calculating DoG using this formula multiple times is much less computationally expensive than applying the original LoG formula each time.

    Visual difference between LoG and DoG graphs. Roughly speaking, DoG can be viewed as a scaled version of LoG.

    Why construct several DoGs withing a single octave?

    Within a single octave, image bluriness gradually increases. The difference between two consecutive DoGs highlights points of interest across scales.

    For instance, a DoG constructed between a pair of consecutive net images (at the bottom of an octave) makes it much easier to detect smaller features. However, for larger features, this is difficult. For that reason, we also compute a DoG for more blurred images (at the top of an octave), where smaller features are not visible, and the algorithm focuses more on larger patches instead.

    SIFT scales appropriately the feature size based on the σ parameter of the DoG layer. Higher values of σ correspond to larger feature sizes.

    Why to construct several octaves?

    It is clear that as blur increases, we can detect larger features. So a natural question arises: why not just use a single octave, iterating from very small blur levels to very high ones? This way, we could detect features of all sizes.

    The motivation for the construction of several octaves lies in two aspects:

    • As blur increases, small details become invisible in the image. So, in terms of efficiency, there is no point in keeping the full image resolution at higher levels of blur. Downsampling reduces the number of pixels by a factor of 4, making processing much faster.

    • Approximating very large Gaussian kernels can accumulate errors, so we shouldn’t use high values of σ. At the same time, downsampling can be roughly thought of as adding blur to the original image, since it also removes fine details. Given that, using higher levels of blur on the full-resolution image can be roughly equivalent to using smaller blur levels on smaller images.

    Therefore, downsampling and octave construction provide significant advantages.

    Why a 3-dimensional window?

    A 3×3 window in a single image can detect local extrema, but there may be too many, especially since we also create multiple scaled versions of the image.

    Adding a third dimension to the window ensures that the detected points of interest are distinctive not only on the 2D plane but also across different scales, making them stable under changes in image zoom.

    If it’s unclear why finding extrema across DoG layers yields points of interest, a useful reminder is that DoG is an approximation of LoG, as defined above. At the same time, we explained in the edge detection article that image edges can be found at extrema after applying the LoG transformation.

    After collecting all potential candidate points, SIFT filters out some of them. The problem is that even if a given point is an extremum, it can still be noise. To keep only the most meaningful ones, SIFT applies a threshold on intensity change to remove low-contrast weak candidates.

    Once interest points are chosen, SIFT tries to construct descriptor representations for them that will allow these features to be matched across different images.

    First of all, detected features across different DoG layers are mapped to circles of varying sizes, where the higher the σ value on the DoG layer, the larger the circle radius. Then, for all pixels in the original image within that circle, gradient directions are computed.

    SIFT then divides the detected region into 4 equal quadrants and constructs a gradient direction distribution for each quadrant.

    Based on the detected extrema point, SIFT draws a circle around the feature neighborhood. It then divides the pixels inside this circle into 4 equal quadrants. For each quadrant, it creates a distribution of gradient directions. These vectors are processed, normalized, and combined to produce a final 128-dimensional feature descriptor.

    4 constructed distributions are then converted into a 128-dimensional vector, which is used as a feature descriptor for the initially detected feature.

    For reference, SIFT provides robust internal mechanisms that allow it to handle situations in which a detected feature has fewer than 128 pixels. SIFT still enables computing a 128-dimensional vector descriptor by taking into account information from neighboring pixels as well.

    A common case in real-world problems is when the same object appears in two images rotated by different amounts. To account for rotation correctly, SIFT also uses additional information about the principal orientation of the gradient, which is simply the most common gradient direction in the distribution. From a rotational perspective, this allows defining the starting point of the object, enabling correct mapping with others and helping avoid false-positive matches. These aspects guarantee the rotational invariance of the SIFT algorithm.

    Comparing SIFT descriptors

    SIFT descriptors are vectors that can be compared numerically to determine how similar they are to each other. The most common use case for descriptor comparison is determining whether the point of interest for which the descriptor is computed is the same across a pair of images.

    L2-distance is a common choice for descriptor comparison:

    L2-distance formula

    The lower the L2-distance, the better the match between two points. If the L2-distance is 0, the match is perfect.

    Another useful metric is histogram intersection:

    Histogram intersection formula

    Here, the formula iterates through each vector component and finds the minimum of two values, which is equivalent to how well a particular aggregated gradient direction is present in both features. In this case, a higher metric value corresponds to better matching results.

    Normally, the same objects across different images are expected to have many matches, making it possible to recognize their identity.

    OpenCV provides an implementation of the SIFT algorithm. To create a SIFT object, the cv2.create_SIFT() method should be called. According to the SIFT documentation, several parameters can be specified:

    • nfeatures: the number of best features to retain. The features are ranked by their scores (measured in SIFT algorithm).

    • nOctaveLayers: the number of layers in each octave. 3 is the value used in the paper.

    • contrastThreshold: the contrast threshold used to filter out weak features in low-contrast regions. The larger the threshold, the less features are produced by the detector.

    • edgeThreshold: the threshold used to filter out edge-like features. The larger the edgeThreshold, the less features are filtered out (more features are retained).

    • sigma: the sigma of the Gaussian applied to the input image at the first octave.

    Apart from the standard algorithm, we can easily visualize detected features on the image using the simple code snippet below.

    import cv2image = cv2.imread('data/input/image.jpg')gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)sift = cv2.SIFT_create()keypoints = sift.detect(gray, None)output = cv2.drawKeypoints(    image,    keypoints,    None,    flags=cv2.DRAW_MATCHES_FLAGS_DRAW_RICH_KEYPOINTS)cv2.imwrite('data/output/image.jpg', output)

    Here is what the result looks like:

    On the left: input image. On the right: detected SIFT features. The lines inside circles, extending from the center to the edge, represent the principal orientation in feature descriptors.

    In reality, for more complex real-life images, the number of detected features can be much higher. Below is another example:

    On the left: input image. On the right: detected SIFT features.

    SIFT has a wide range of applications. Let’s look at them.

    Image matching

    As mentioned before, feature descriptors can be used for image matching. Let’s look at one example using the following image pair:

    A pair of input images. Both images contain the same objects but are composed in slightly different ways, with some objects placed in different positions, at different scales, or with different rotation angles.

    First, we will read a pair of images. Remember that before feeding them to SIFT, they must be converted to grayscale.

    import cv2IMAGE_ONE_PATH = "data/input/image_1.jpg"IMAGE_TWO_PATH = "data/input/image_2.jpg"OUTPUT_PATH = "data/output/matches.png"image_one = cv2.imread(IMAGE_ONE_PATH)image_two = cv2.imread(IMAGE_TWO_PATH)gray_one = cv2.cvtColor(image_one, cv2.COLOR_BGR2GRAY)gray_two = cv2.cvtColor(image_two, cv2.COLOR_BGR2GRAY)

    We then compute descriptors for each image.

    sift = cv2.SIFT_create()keypoints_one, descriptors_one = sift.detectAndCompute(gray_one, None)keypoints_two, descriptors_two = sift.detectAndCompute(gray_two, None)

    Next, we’ll utilize BFMMatcher, or Brute-Force Matcher. This algorithm compares every descriptor in the first image with every descriptor in the second image, identifying the closest pairs based on the chosen distance measure. In our code, we’re using the L2-distance.

    By calling the knnMatch() method, we pass all descriptors from both images and set k = 2, which specifies how many of the top k closest matches are returned for each descriptor.

    matcher = cv2.BFMatcher(cv2.NORM_L2)knn_matches = matcher.knnMatch(descriptors_one, descriptors_two, k=2)

    The goal of setting the parameter k to a value greater than 1 is to remove less relevant matches by using Lowe’s ratio test.

    The test consists of determining how good the best match is compared with the second-best match.

    RATIO = 0.75MAX_MATCHES = 50

    For instance, in the code below, we filter only good matches where the distance to the best match is less than the distance to the second-best match, scaled by RATIO = 0.75.

    As it turns out, we can still end up with too many matches, so we keep only the best MAX_MATCHES = 50.

    good_matches = [best_match for best_match, second_best_match in knn_matches if best_match.distance < RATIO * second_best_match.distance]good_matches.sort(key=lambda m: m.distance)good_matches = good_matches[:MAX_MATCHES]

    Finally, we can draw matches using the cv2.drawMatches() function.

    image_one_padded = cv2.copyMakeBorder(    image_one, 0, 0, 0,    20, cv2.BORDER_CONSTANT, value=(255, 255, 255),) # adds a small white marginmatched = cv2.drawMatches(    image_one_padded, keypoints_one,    image_two, keypoints_two,    good_matches, None,    flags=cv2.DrawMatchesFlags_NOT_DRAW_SINGLE_POINTS,)cv2.imwrite(OUTPUT_PATH, matched)

    Here is the result:

    Lines showing matched features by SIFT in both images.

    As we can see, SIFT did its job very well! It correctly matched the main objects in both scenes. A very interesting observation is that SIFT successfully preserved scale and rotation invariance!

    For example, we can see that the plane was scaled and rotated differently in each scene. Despite this, SIFT produced very similar descriptors for each plane feature.

    Object detection

    Another SIFT application is object detection. With an object template, we can perform image matching in the same manner as above to search for that object in an image.

    Recognized objects in the scene using SIFT from the templates on the right.

    A great aspect of SIFT is that it tends to be robust against occlusions. If a part of an object is overlaid by another object, SIFT can still detect the visible features and match them successfully.

    Once feature matching is complete, additional postprocessing techniques can be applied to extract the object’s contour and determine its precise location in the image.

    It is worth noting that SIFT can sometimes produce false-positive matches, as shown in the image above. We can clearly see a red line connecting the airplane’s left wing in the scene on the left to its endpoint in the object template on the right, where it matches a point on the right wing.

    Such situations can occur from time to time, and in most cases, they do not have a strongly negative impact. Depending on the task, postprocessing algorithms (e.g., RANSAC) can eliminate false-positive matches if there are not too many.

    Image stitching

    Image stitching is the task of merging photographic images taken from a single viewpoint that have overlapping regions into a single high-resolution image (a panorama). Image stitching can be elegantly solved with feature matching, perspective warping, and geometric transformations.

    To do that, it is necessary to understand homography, which we will cover in one of the next articles.

    3D-reconstruction?

    While SIFT works very well for flat and 2D objects, it is unfortunately not suitable alone for matching 3D objects.

    However, SIFT is used as one of the core steps in other 3D reconstruction algorithms (e.g., COLMAP). It allows matching points across images, from which the whole 3D scene is then constructed.

    SIFT is a highly versatile algorithm for feature matching, notable for its capacity to match features while maintaining rotation and scale invariance.

    As we saw, SIFT can solve a wide range of problems in computer vision. Image matching, object detection, and image stitching are among the most popular SIFT applications. In more sophisticated problems, SIFT is often used as a strong backbone for feature matching, which is then processed differently depending on the problem itself.

    All images unless otherwise noted are by the author.

    Algorithm Computer feature Invariant Scale SIFT Transform vision
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Previous ArticleCan Safeworld convince people that gen AI robots won’t hurt them?
    Next Article AI glasses face their first major government crackdown
    • Website

    Related Posts

    AI Tools

    How to Build a Cheap, Yet Reliable Model Router With Jev

    AI Tools

    How to Use a Free AI Generator to Ship a Full Campaign in 90 Minutes

    AI Tools

    How to Govern AI Agents

    Add A Comment
    Leave A Reply Cancel Reply

    Top Posts

    Qwen3.8 27B addition in words

    0 Views

    Connecting AI agents to enterprise knowledge

    0 Views

    OpenAI launches visual ads that appear alongside image generation results

    0 Views
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram
    Latest Reviews
    AI Tutorials

    Quantization from the ground up

    AI Tools

    David Sacks is done as AI czar — here’s what he’s doing instead

    AI Reviews

    Judge sides with Anthropic to temporarily block the Pentagon’s ban

    Subscribe to Updates

    Get the latest tech news from FooBar about tech, design and biz.

    Most Popular

    Qwen3.8 27B addition in words

    0 Views

    Connecting AI agents to enterprise knowledge

    0 Views

    OpenAI launches visual ads that appear alongside image generation results

    0 Views
    Our Picks

    Quantization from the ground up

    David Sacks is done as AI czar — here’s what he’s doing instead

    Judge sides with Anthropic to temporarily block the Pentagon’s ban

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Facebook X (Twitter) Instagram Pinterest
    • About Us
    • Contact Us
    • Terms & Conditions
    • Privacy Policy
    • Disclaimer

    © 2026 ainewstoday.co. All rights reserved. Designed by DD.

    Type above and press Enter to search. Press Esc to cancel.