Computer Vision: LIDAR, Time-of-flight sensing and structured light scanning

It's necessary to get 3D factors when filming sets and movie sets, and I want to discuss different ways to get 3D data using laser ranging. There's LIDAR, and Structured Light Scanning, as many. LIDAR is light detection and ranging, and pulses come out of a window, and hit objects of a scene, measuring the time it took from the object to come to the scene and back. 

It tells us how far away a surface is.



You're using a physical object in order to judge the distance. We use a type of laser or structured light scanner if we really want to get accurate 3d imaging.

We get a sensor and probes and object in the scene and the pulse comes back to the sensore and we get a distance map dependent of θ and Φ, and this can be interpreted as a depth image or a range image.





We're also familiar with Google Cards, with a laser scanner on the top of the car. The surface is made of a column of points.


We want to get the physical surface of the building that I'm scanning, and we're getting a bit wooden ball inside the building whenever I'm scanning part of the building.

You can't get 3D points on any surface you can't see with just the naked eye. Unfortunately these shadow spread more the further away and object is. I want to get the physical surface of where the object is.  The laser goes right through the glass, so we're gonna have a problem. If we have an object, we rely on the physical surface to return some laser intensity back to you.

Most laser intensity will reflect onto a surface back to the surface. Sometimes you may miss surfaces that are parallel to laser rays due to the "grazing angle effect". The further away the object is, the more that laser spot will expand. If I have a red laser then I have a green object, a white object will reflect more light from the camera than the green object, giving a good clue on how well I scanned the object. The best case scenario is I'm scanning objects that are perpendicular to the new angle. If I wanted to scan a beautiful new black, shiny car this would be problematic, because the surface will absorb light and act like a mirror, and unfortunately they will bounce back off and never come to a scanner. They bounce back and return to the object.




There's a very simple equation called pulse-based LIDAR, and I want to figure out a specific distance d. I have a pulse that travelled a distance of 2D. We take T time units which means 2d/t = c, where c is the speed of light where d = 1/2 ct, which is the key behind the time of light scanner. So to get a 5mm accuracy, the time of light is 33 picoseconds, which the same thing as sensing timing accuracy. The receiving accuracy need to be extremely accurate. Velodyne is pulse-based. You have to shell out some serious cash to put that together, for a very high accuracy receiver. The return light intensity might by much smaller, which means that the sensor needs to be extremely sensitive in terms of photons, but you can mitigate noise in the system, by sending out multiple pulses and average them, meaning that the variants of those measurements go down. The resolution of the scanner is mm accuracy. 

The idea behind phase-based LIDAR is different, is whenever we send a pulse we say there is a sine wave with a very high frequency. We can shape the envelope of a sin wave with a much slower frequency wave, and I can modulate this frequency much more slower. 

We can imaging we have a sensor and object and we are sending out a portion of a slowly moving sine wave that hits the surface and comes back, and I measure the different in the phase of what I received. There hasn't been much time for the slow frequency sine wave. What I look at is psi(Ψ), which is the phase shift of the modulator signal. What I have is ω is the frequency of signal, where Ψ = ωt where therefore we calculate d = cΨ/(2ω), which makes this technique faster than the pulse-based method.




 The downside is that there is an ambiguity with an object that happens to be one period or half a period away from the scanner, we get one period. But if we have an object, if the phase shift is the same, we can't distinguish things property.

The most physical process is "I know my sensor is unambiguous within a certain range." Typically, I am not against any ambiguity, so in general I don't really have a problem with this.


We take one of the LIDAR scanners, then I can acquire distances to point, and we classify 3D pointers into different classifications, and this helps visual effects people look at what to do at the pipelines. If all I got was the 3D return, it is very difficult. 

We want to simulate scenes like crash a car further without actually crashing it.

Velodyne scanner acquire rings of data the car says where a car is, where a person is.

1. I need fairly accurate inertial measurements and where the car is. 





We can also preserve cultural heritage with 3D scanners.

A different alternative to LIDAR is the time-of-flight camera. 


The Velodyne scanner does live scanning, but we want to do live realtime 3D scanning. Compared to a general camera, these things are generally pretty low resolution, but the advantage is they can do these scanning images.

The 3rd type is the Flash LIDAR or the TOF camera, and these are usually fast (30 Hz frame weight) and they are usually low resolution, and this was a SR 4000 sensor from MESA imaging as a 176 x 144 image, and I would say that it's noisy, meaning the measurement is accurate and we can see some visible noise, and usually it's very easy to set up, suddenly getting some real time 3D depth. The downside is that there can be a whole bunch of missing stuff, and we can get no responses on certain regions of an image. 

There's time of flight for detecting occupancy of a smart room.



  We got some small sensors. One advantage of them is that time-of-flight makes us feel good, and we don't want to feel like our rooms are modifying us, and we don't want to feel like someone is hacking the cameras. We should have chunk applications and just determine what people are doing.

The sensors themselves are pretty noisy, but they are good enough to track various people as they move around the room. Our next phase is to use the real-time locations to develop control algorithms for the lights in the room.

The most notable current use of this is for the Xbox One Kinect version 2.

 

Since it came out, we can take a beta kit and show some results. We can keep performing this to figure out what the resolution of an image is.

We can see characteristic noise around the edges of a particular image, and it is purpose is good in terms of consumer game-control.

Kinect can likely be a result of a decision-tree classifier. It's probably trained more or less on expected poses as opposed to actual poses, so it's pretty neat.


The next lecture is about structured light scanning. Last time we talked about LIDAR scanning, and we use LIDAR if you scan something as big as a car, but structured light scanning is what we do if we want to scan something on a tabletop, or an actor in costume, or other various objects. 







Structured light is looking how a light deforms and use those deformities to determine 3D shape. 

When the stripe is falling out a surface that is smoothly varying, you get smoothness, otherwise you get a kink in the stripe and when the stripe falls out 2 different surfaces, you get a discontinuity in the stripe, and this is all basically just a matter of discalibration.


We have a homography between the points in the image plane and the points in the laser plane.

Structured light Scanning:

The first method we want to talk about is with a laser stripe, with the idea being to compute the projective transformation between the plane of the laser in 3d and the image plane.

The first idea is with a laser stripe and we want to compute the projective transformation between the plane of the laser in 3D and the image plane. Let's say there's a calibration object with planar faces. When the camera sees that object we say "That is equivalent to seeing the checkerboard in a bunch of configurations. If all the faces are identical, we see what we need to calibrate the images. I know all things on the red stripes has to hit the points in the 3-dimensional plane.


This only gives the laser plan between the position of the laser and the position of the camera. I have to calibrate the system for every possible location of the laser plane, which can overcomplicate things.

A romer arm, knows exactly where it has been put in 3D space, and scans everywhere, possible to stich and calibrate images? This arm has to be mounted very rigidly, and it should be a table with high amount of friction. It scans different materials, and a very common use is for things like CAD modelling.



This is why scanners usually rotate with many platforms, and has a slice as a stripe move up and down, and we have the cameras inside of them, and are super calibrated against each other.



We wanted to scan the statue to discern the chisel strokes of Michelangelo. They do this for putting data to a file, but it's not burdensome to be scanned at all. You would get a decent level of detail for these kinds of scans, because you can localize these lasers in detail and get high quality images.



The projected lasers still has to be a good distance from the position of the camera. We can calibrate cameras and say "I know where these cameras are in the cameras are in the 3D space, and this makes me able to compute the epipolar geometry for this space."

I know where I am for every position and for every position I can triangulate that curve in 3D space. The structured light only put these coordinates on the scene.

Instead of trying to see an image, I can slowly move my stripe across the surface of an object, and every stripe is representative of a Gaussian Object. I can plot the intensity and accurately locate where the center of a stripe is, and I can look at the intensity of the stripe as it passes across the pixel. This is called space-time analysis.



We don't care what we're predicting, we just want to get matches among the stripes.

We move camera and we use that to very accurately triangulate where a stripe is.

Why don't we just get n slices of an object all at once? It turns out to be disambiguate which stripe is and we accidentally lose trips and get them mixed, and we can't tell which stripe is which. We need to code the stripes to tell which stripe is which.

Time multiplexing is instead of looking at one image, we shine a bunch of patterns on an object and we ask how do the time positions of these objects uniquely give me a position in space? 

I only need to show n patterns to get 2^n stripes for patterns, but we are eventually limited on how finely I can project the stripes.

These patterns that are uniquely compatible are called graycodes, which are alternating patterns of black and white are shifted by a period, and each stripe and its neighbor only differs by one bit of difference.



For every pattern, I show the pattern and I show its opposite. We can try all sorts of things for all the time in the world to scan. It's very difficult for someone to scale as we scan n light patterns. I can also encode the position of a stripe by looking at what happened on either side of it. We look in the line in between several stripes. We see the difference of a certain number of stripes, and we lose some patterns thinking about the adjacent strips instead of just the stripe.

We shine unique patterns of color onto the object, and We show a colorful pattern and the idea is that every unique stripes are unique, and the alteration of colors tells me the stripe index. The idea is called a de Bruijn sequence where there is no repeat among a certain set of stripes. 

000001010100101110111 and a sequence of symbols will be converted to color stripe codes.

We  memorize the sequence of the cards. We shine something of the object, and then we take the things on the pattern and match that up using dynamic programming, and we subsequently know the order of the colors.


We have issues if stripes miss or switch places. We talk about how things correspond. Again, we may have to do multiple steps of dynamic programming in order to disambiguate this.

What are the advantages of these color stripes? We can show this one pattern at once!

We have a moving object and we can reconstruct the shape in real time, and shine one pattern of light and get 3d immediately. It's easy this way because we can scan features that are moving.

We can also plot a cosine wave, with 3 continuous numbers and we can also extract the 3D position of the point, and for this we need to project 3 points onto the object.




The kinect was projecting a pseudo random 2d dot pattern onto the world, and these dots are not random because every point in space has a unique neighborhood of dots. If I see this dot and look at all my neighbors, I know where my projection was in the case of the projector. Scanners have both a camera and a projector of light, but you can't see what the pattern is, because it is a white light at a very high stroke rate. 

The histogram on the left tells what the "sweet box" from scanner, and the images are real-time sensor of the scanner.


Things look pretty good once you texture map it. Getting all the nooks and crannies are pritty tricky like getting under the chin, etc. We can do a watertype mesh and stuff. We have scans for visual effects, movies, and toys, and now action figures look good because they come from the 3D scans that come on set.


Figures will be very carefully 3D-sculpted and projected later on and there's a lot of digital cleanup work to make the right quality scan for the job at hand.


These things are used for toys and the large displays from comicon.

Thanks for reading!


Comments

Popular Posts