Stanford AA274A Principles of Robotic Autonomy | Autumn 2019 | Sensors & Intro to CV
Stanford Online
Sensor goals 0:05
The lecture shifts to robot perception. That means gathering information about the world, pulling out useful features, and using them to localize the robot and choose the next action. It sets up the main sensor themes, then turns quickly to cameras, which will matter most in the course and on the class platform you will use.
Sensor types and errors 4:32
Sensors are first split into proprioceptive and exteroceptive, then into passive and active. A camera is passive because it measures light, while a laser rangefinder or a camera with a flash is active because it sends energy out first. The lecture then defines design specs like dynamic range, resolution, linearity, and bandwidth, and field specs like sensitivity, cross sensitivity, error, accuracy, and precision. It ends with systematic errors that calibration can fix and random errors that must be modeled and carried through the perception pipeline.
Heading sensors 20:02
Heading sensors tell you which way the robot faces relative to a reference frame. A gyroscope uses a spinning wheel or disc whose axis stays stable, so it can infer heading, while a compass is an external sensor for the same job. Gyroscopes are common in IMUs, but small amounts of friction cause drift, so you need other sensors from time to time to correct the heading.
Acceleration and ranging 23:01
Accelerometers use a spring, damper, and proof mass to measure force. If the force is steady, the mass settles, and the final displacement gives the applied acceleration. Modern versions are tiny micromechanical parts, but they are fragile. The lecture then turns to IMUs, which combine gyroscope and accelerometer readings to estimate motion, and to active ranging sensors. Time-of-flight sensors measure how long a wave takes to return, as with lidar, while structured-light sensors project a pattern and read its distortion. Lidar can reach farther, but it is more expensive and must stay safe for the human eye.
Other sensors 38:02
Radar is now being designed for robot autonomy, not just for airplanes and other distant targets. It is being adapted for structured places like cities. New sensors are also appearing, including tactile sensors, artificial skins that measure pressure and temperature, and neuromorphic cameras that work more like the human eye.
Pinhole camera model 40:00
The focus shifts to cameras and basic computer vision. A camera captures light through a lens and turns it into a digital image. The simplest model is the pinhole camera, where a tiny hole blocks most light so that each point in the scene maps to one point on the image surface. That gives a sharp, faithful image, though it is inverted.
Perspective projection 47:01
The camera has a focal length, which is the distance from the pinhole to the image plane. A virtual image plane is often used because it makes the math easier and avoids the inversion. With a camera frame centered at the pinhole, a 3D point with coordinates X, Y, Z projects to image coordinates x and y by perspective projection. The result is x = fX/Z and y = fY/Z when the focal length f is known.
Pinhole trade-off 57:30
A smaller pinhole makes the image crisper, as shown by the sequence from 2 mm down to 0.35 mm. But less light reaches the sensor, so the image becomes dark. If the hole gets too small, diffraction also becomes a problem. The fix is to use a lens, which focuses light by refraction instead of relying on a tiny opening.
Lens and focus 59:34
A thin lens gathers many rays from a point and brings them to one point on the image surface. That gives the same one-to-one mapping as a pinhole camera, but with a much brighter image. Objects at the wrong distance can blur because a lens focuses best at a particular range. For normal use, that is acceptable, and the camera can be treated like a pinhole model.
Thin lens relation 1:04:01
Using similar triangles, the notes lead to the thin lens equation, which ties together focal length, object distance, and image distance. A faraway object focuses at the focal length, so objects at infinity land there. This is why a camera with a lens can be modeled the same way as a pinhole camera, as long as you use the focal length in place of the pinhole distance.
From world to pixels 1:12:31
The next step is to map a world point into camera coordinates, then into the image plane, and then into pixel coordinates. The direct camera-to-image step uses the same formulas already shown: x equals f times Xc over Zc, and y equals f times Yc over Zc. The remaining steps will wait for next time, because the mapping has a division by Zc and is nonlinear. Homogeneous coordinates will be used to make it linear, which will support camera calibration and later scene reconstruction.
AI-generated summary. It can be wrong or incomplete - check anything that matters against the original.

