Why camera calibration is so important in computer vision
opencv.ai
opencv.ai
Uncertainty propagation is very difficult to use for vision, and largely just modelling errors in vision as the error distributions, e.g. for anything observed or reconstructed from images are either subpixel accurate, or too non linear.
Another problem I run into is that the transforms are ill-defined just outside of the screen. That means that if I want to draw e.g. a line in world coordinates onto the image from a camera, then I often get garbage if the line starts or ends outside the image (even if I divide the line into many segments).
Yeah thats a common problem if the calibration failed. It could also be that you are not cropping to what is in front of the camera, but if its really weird, its most likely the former.
So the default, and probably most of the camera models in opencv requires a monotonic change lenses and bijective imaging. The former is common unless the lense has defects on the surface, and the latter is practically a physical constraint. The problem is that these constraints are difficult to add to the estimator, so they didnt. Meaning it will find a solution where they are not satisfied. If say the bijectiveness is not satisfied a bit outside the image, but stil valid accounting for infront and float accuracy, then that would absolutely account for the problem you describe. Is pretty obvious if you consider the function what the problem is, just hard to add the constraint in opencvs estimator.
The solution is 1, verify the result is satisfied after estimation, 2 make sure you make the parameters are as observable as possible during calibration. This means spread out in the image, evenly distributed, and all the way to the edges. Also make verify it has not rotated one or more of the detected chessboards upside down, or 90 deg sideways. Finally, because it becomes harder and harder to avoid this problem with more parameters, always start with 1, then try 2, then the two variations of 3, and so on. More parameters always fit better, so use an appropriate test.
Can you be more specific on the second point? Once you have the intrinsics, it's trivial to project them to an (undistorted) image.
Opencv calibration is hard to get right if you dont know what you are doing, and at first glance it will look right despite being irrecoverably bad. I had to edit the docs and tutorials because the examples they had were of FAILED calibrations. Visualizing the resulting distortion field is critical to understand if the calibration succeeded, and that s what the better tools provide. That said, if you dont know you are looking at, they likely dont help, and the golden rule is if someone calibrated something and said it was easy, they didnt. If someone has a pdh in camera calibration said they calibratied them but that they havent used them for sfm but they plan to soon, odds are 20% its right.
Sfm tests the calibration to the extreme and even single pixel calibration error will be highlighted as correlated reprojection error vector in most images. And that is assuming the bad calibration does not cause the system to fail outright. There is also the inbetween case where the system is only able to use small parts of the image.
The distortion field often visibly tells you if the estimation failed. It should be smooth, and monotonic, and you can draw it not just for where there are pixels but further and in higher resolution, and for regular lenses almost always highly symmetric. Looking at the distortion field you can see alot of problems that you could not otherwise. The most common problem is the monotonicity, since that constraint is very difficult to add to general optimizers. Since this means the distortion goes backwards, they are visible as sharp edges in the distortion field.
There are applications that require pulling almost magical levels of signal from very small numbers of pixels, and for those extremely accurate results you really need to drill down on the camera model in particular. Even modelling the environment or atmosphere to get accurate results.
For those types of applications opencv is absolutely not the right choice. It just so happens that dima has an accessible software kit that is more for the latter case. But all the better because it's perfectly fine to use this tool to calibrate any camera and then go use it with any other CV pipeline. His pain your gain.
Such algorithms always create a graph over the images, but if you mean the graphslam graph filter methods, those are substantially subpar compared to classic feature based methods as well as the more modern, dense and semidense methods.
For localization work (indoor or outdoor), in which you mostly want to close loops and track 3D pose of features using measurements from, say 100s of m or less, (or perhaps scene recognition using effectively flat "very far" images), calibration using opencv is probably highly performant. It's standard, and I've certainly seen much success using it before feeding into regular gtsam etc. There are some unspoken assumptions about localization that don't translate to, say, tracking necessarily. Generally they are: Many observations, close-ish range (relative to stereo offset), mostly-correctly-pointed camera rigs (e.g., forward on a car), or perhaps assumptions about density of features, correlation across images, existence of dense prior maps, etc.
I believe the use case for something like mrcal is to improve calibration for cameras and applications that don't fit this model well. In particular, you may need to track a target at extreme range, using a pixel or two in the corner of your image, with a particularly wide FOV. These specific use cases, mentioned in top level comment, do require additional care, especially in calibration. Thus mrcal.
It just so happens that out of a domain where the very particular details of calibration matter, you may find a tool that helps with all calibration and I think that's the point in bringing up mrcal on a thread discussing calibration in general.
I think that's as precise and non-controversial as I can be.
I've been working in AI/ML and some CV projects over the last decade.
I've seen far too many cases where the "algorithms" and modeling teams had no concern or even a concept for how the input systems for data that would be used to train models, and later inference, mattered to the quality of outcomes.
In CV and computational photography cases, there was little concern or understanding for how photography and imaging actually works, nor consideration for how variances in hardware, configuration, or the capture pipeline in general might affect those models when doing training data capture in parallel. Adding another layer, consideration for how variances between the capture pipeline and actual inference pipelines might need to be accounted for when designing the overall system and how to approach data collection and curation for training + evaluation, as well as to build-in robustness. (similar thoughts apply to concepts like bias and fairness in models)
Example: training vision models using one set of imaging hardware and configurations while applying those models on very different imaging hardware with different characteristics.
To summarize the above, not caring about calibration in CV is like not caring about how variances in feature extraction/embedding generation will affect the overall quality of your results.
But then it all crashed. Somewhere along the line we lost something. My last job (not Shield), I watched the estimates for a truck jump in and out of existence at 50m range with 2m baseline. Its orientation flickered 30degrees or more. But because they had used an EKF somewhere, they claimed it was the best they could do and anything else was inventing signal. I haven't dug into state estimation in the last five years but it's not gotten better if this is what new grads believe.
Many of the libraries are not well maintained. I needed a basic homography estimation recently and asked a minion to try opencv for it before anything more advanced. He got it working, but the best inlier ratio for every feature he tried with default parameters was 3%. The images were offset by 30 pixels left right, and less than a pixel in warp and taken with different exposure time. So he used sift...
He argued he got it working and that was as good as keypoint matching got... If I hadnt happened to be the guy who needed it one step up, it could have just propagated, someone adding a shitty ekf to make it smooth, then buried in layers of heuristics and api.
Many "AI/ML" teams have been doing more alchemy than science.
One memorable team "revelation" was when they figured out the hard way that many objects and human subjects, in particular, register completely differently based on the imaging spectrum used (e.g. human visible light, versus near infrared wavelengths).
People familiar with the "old ways" probably would have considered this to be obvious, but it seems that the "new ways" are often correlated with an absence of domain knowledge related to what the "new ways" are being applied to. This isn't a new story, and history seems to repeat itself.
The lense model is the parametrized approximation of the lense function. Picking the right one matters for ideal accuracy.
The estimator takes observations which constrain the parameters of the lens function, and find the parameters.
The framework helps you create those observations.
I'm pretty confident that all cameras used in space exploration missions have been thoroughly characterized on Earth labs before launch.
For example here is paper describing the calibration activities for Curiosity rover: https://agupubs.onlinelibrary.wiley.com/doi/10.1002/2016EA00...
1. camera calibration 2. lighting
Lighting is the easy part
> We don't know where the camera will be. For example, we ask players to use the cameras of their mobile devices to explore the size and proportions of a room — to work with the augmented reality helmet. We don't know anything about the rooms, and we can't ask users to use a pattern.
Are there automatic techniques to tackle this problem?
What makes it harder is that older phones have insufficient processing power, whereas newer phones have too many cameras. The phones also do post processing in undocumented ways that are impossible to disable. There is also rolling shutter. This in turn can be simplified down again by just subsampling the images to 480p and estimating camera parameters per image.
In general a good starting point for such a problem would be to implement https://grail.cs.washington.edu/rome/ on your own.
Its not particularily hard, and there is no difference between reconstructing from uncalibrated cameras and reconstructing from many images from one camera with unknown calibration. Then you can check if the calibration is unchanged over time, and if it is, add that as a constraint. So far it will still take half an hour to run, but now you can add in tracking exploiting motion prediction and imu, which in turn will give you the real time speed.
Extrinsic calibration (where is my camera in 3D space, and how is it oriented?) being different from intrinsic calibration (how do the pixels in my image map to ray directions relative to the front nodal point of the lens?).
https://www.semanticscholar.org/paper/Body-relative-navigati...
Auto calibration/calibration in the loop is not really "no calibration".
OCR for distorted text is a solved problem. Relatively simple math based on normalizing text size and shape can be used to calculate how to reverse distortion.
Similarly, maybe the next frontier in 3D vision is employing large datasets to train a black-box AI without us having to understand and get the math right.
Ergo, I would not bet against teams that understand that objective technical field; AI is not magic, and knowing the fundamentals will always help. Better models, smaller models, faster models; the less you leave for the model to solve in latent space the better you’ll do, I think.
You can argue the exact opposite: the more you leave to the model to learn by itself, the more likely you are to find a solution that was not accessible to humans and their limited feature engineering capability.
I mean yeah, but thats because text is a fairly predictable, simple thing to extract. It has sharp edges, limited number of lines, and is neatly spaced relative to each glyph. Its also mostly high contrast.
Structure from motion however has none of those luxuries. for "AI" to overcome that, it would require a model that has an innate understanding of the scale, shape and orientation of most objects in the world. Not only that, it'd also need to have an innate understanding of lens distortion to work out if the object is bent because of the image, or the object, or some other effect.
I look forward to you releasing your model that does all this.
Not saying it's easy, but solving it in software is O(1) and solving it by calibrating each device is O(N), and allows you to be resilient to things like temperature deformations and other nasty things that can happen in the field.
I can't find the reference any more, but one particularly impressive result was a four-wheeled robot that could tolerate the loss of a wheel, distortions of the entire chassis, and multiple faulty sensors all at the same time!
For critical imaging, you still want to calibrate, but the gap between explicit calibration and "implicit calibration" (just using your subject matter) is definitely getting smaller, and quickly.
The browser doesn't let you collect all this sensor data.
Mobile Safari doesn't tell you the known camera calibration.
The Apple Vision Pro doesn't give you any data at all.
Nonetheless, 8thwall's XR8 takes seconds to calibrate absolute scale. It has to do device lookup for approximate calibrations for your iPhone model.
If I bought a car and 5 min into driving it tells me that the camera system is defective you are not gonna have a happy customer.
Take the other example then: "pre flight calibration when one first takes it out of the box".
If there was an issue with the camera system it could be caught at that time.
It even has an option in the user Service menu to recalibrate them.
But I didn't know a coin was used either.
That said, there's a long tradition of using circles and checkerboards that goes back to early BBC TV broadcasts transmitting a tuning pattern.
The coin used is a circle, likely has "interesting" patterns, and perhaps even has raised and gouged features that the stereo paired cameras can height resolve (and calibrate against for when examining rocks).
"Cameras, by design, also introduce a level of distortion": distortion is an artifact that is usually minimized in lens design if you are trying to create a rectilinear lens, though you are typically forced to trade it off versus other design criteria (like resolving power, or optical complexity). If you're trying to design a rectilinear lens, you explicitly try to remove any distortion; for "normal" focal lengths (focal length ~= image circle diagonal) you can often get vanishingly low distortion, even with relatively simple lenses. (There was a 19th century lens design that achieved "zero" distortion with just a few elements, and a precisely-chosen aperture position; the name escapes me at the moment.) If you are not designing a rectilinear lens, there are other lens mappings, in which case it's not really proper to describe the effect as distortion, though in technical literature it's often still described this way. (For further reading, check out F-theta vs. F-tan-theta lenses: https://en.wikipedia.org/wiki/Fisheye_lens#Focal_length; https://www.thorlabs.com/newgrouppage9.cfm?objectgroup_id=10....)
"... the proportions of objects closer to the edges of the frame are distorted." This is technically incorrect; objects subtend the same number of pixels in an f*theta fisheye image whether they're at the center of the frame or at the edge, for a given camera-subject distance. The apparent distortion exists because you're viewing it further away than the focal length implies, so the viewer at a "normal" viewing distance is applying a transform to the image that is not distance-preserving. If you get extremely close to the image, and if it were displayed on a curved surface (so the image's angular extent to the viewer matched the camera's FOV), you'd see no distortion. Also, we think of arranging subject matter for, e.g., a portrait, along a plane that's orthogonal to the camera's axis -- but then people closer to the edge are farther from the camera, and would naturally appear smaller. It's a weird artifact of rectilinear lenses that they're actually magnified in the resulting image so people standing in a line are all equal size; to get the same effect with a fisheye lens, you just have people arranged in an arc centered on the camera. (Most portraits are not shot on fisheye lenses though, for very good reasons, so this is more of an edge case.)
Related: "there are no straight lines" -- if you're using a camera to make an image, you're operating in projective space. Whether a line is straight in 3D space, and is then straight in the projective space, is a function of your projection. Rectilinear lenses -- for which the relation f*tan(theta) applies -- straight lines remain straight. But for most (all?) other mappings they do not; for example, f*theta lenses preserve angular extent: an object that subtends a particular angle from the camera's perspective will occupy the same number of pixels regardless of where they appear in the image. This is arguably superior for many computer vision tasks.
... I need to stop here, because this is already getting tedious. For anyone interested in developing a deep intuition for lenses and imaging, I would recommend playing around with a view camera, understanding the Scheimpflug relationship (https://en.wikipedia.org/wiki/Scheimpflug_principle), and really thinking through imaging from first principles if you want to understand these things more deeply.
Or just, you know, use Colmap and don't look back. That works too.
Calibration in this context is essentially the task of finding the optimal parameters of some (usually nonlinear) function (u,v)=f(x,y) that remaps positions in the original image frame to a rectified frame, where all straight lines in the world appear straight in the image. Technically, a skewed and squashed image would also fulfill those requirements, too. But this is a customer-oriented blog post, to give someone enough of an understanding to convince them of the importance of calibration, it's not a rigorous technical paper, so I actually think it's fine to skip/simplify some details.
Now, I may be biased, because I work in imaging-for-humans, and I've had many a conversations with engineers about why a particular simplification doesn't work for, e.g., filmmakers, but I think that even for purely technical disciplines, understanding the assumptions that go into the pinhole model can be useful. At the margins. Which sometimes matter.
This is often elided by camera systems, that apply a gain to peripheral pixels to correct for this phenomenon. If you understand imaging, you will expect that, and understand why, for example, your wide angle lens displays a lower signal-to-noise ratio for a given illumination value than you might otherwise expect at the edges of the image.
This is a really specific example, but there are dozens. Imaging is its own deep, technical field that is abstracted, and occasionally obscured, by the pinhole model.
That rectilinear projection is indeed the most popular choice, but all projections involve tradeoffs as the FOV gets bigger, in the same way that all planar cartographic projections involve tradeoffs as the depicted region gets bigger. For example, the magnification of objects at the edges of a rectilinear projection gets extreme as the FOV approaches 180 degrees, and the projection stops existing entirely at or beyond that. That magnification is sometimes called "perspective distortion", even though it's inherent to the rectilinear projection.
Wide-angle lenses or multi-lens arrays will often deliberately choose f*theta instead, to avoid that "perspective distortion" or support FOV >= 180 degrees. Other projections (e.g. equirectangular) are also used, especially for stuff like panoramas and VR. The concept of distortion is meaningful only with respect to a desired baseline, which is often but not always rectilinear.
IMO this is the correct definition of distortion. However, as the parent comment said:
> If you are not designing a rectilinear lens, there are other lens mappings, in which case it's not really proper to describe the effect as distortion, though in technical literature it's often still described this way.
I think many people confuse mapping and distortion. When a fisheye lens is used, it's often seen as "heavy distortion". But a more accurate way should be to say that it's a different mapping/projection, and the _distortion_ a calibration measures is the difference of this ideal projection (ftheta (e.g. Kannala Brandt) rather than fthan*theta (pinhole) ), and the actual image. This can be a minuscule amount. This means that, "undistorting" a fisheye image doesn't give you a rectilinear image, but still a fisheye image. You can of course decide to map the undistorted fisheye image to a rectilinear one, but that's conceptually a different operation than (un)distortion.