How to design a SLAM system

The SLAM algorithm based on traditional computer vision geometry and state estimation has been mature. There are lots of open source SLAM architecture, and different commercial products have their own implementation. Every SLAM implementation has its pros and cons. This post is to provide the high level guidance for how to design a practical SLAM system for products.

Is SLAM needed?

The first question to ask is whether SLAM is really needed. SLAM algorithm has 2 parts, localization and mapping. It could provide very accurate 6DoF pose relative to the starting point, and build 3D point cloud of the nearby environment without any prior info. However, SLAM system usually cost lots of compute resources and memory.

According to the requirements and available resources, SLAM may not needed. e.g. If the system just want to achieve a coarse outdoor localization, GNSS might be enough. If there are some outside-in tracking systems like the base stations of HTC Vive, SLAM is not required either.

A SLAM system is needed if

  • Long period, realtime, accurate 6DoF pose needed
  • And, not enough knowledge of the environment
  • And, no outside-in tracking support or not enough

Targeting Scenario

Based on different targeting scenarios, the design and implementation of a SLAM system could vary. Taking AR/XR use cases as an examples -

For world-locked rendering in small/stationary regions, it is sufficient to only have the frontend of SLAM, e.g. open-looped VIO. Tracking in a small region won’t drift when maintaining a good local map, and the map could be threw away after each session. Therefore the loop closure detections and complex map management won’t be needed, while the tracking latency and smooth turn to be important factors.

For world-locked rendering in large regions targeting virtual object persistence, or cross-session/cross-device object persistence, the full SLAM system is needed, including loop closure/place recognition and map management. The loop closure could helps to resolve the pose drift over closed-loop long trajectory, and map management is the key to anchor virtual object into the physical space. The map data needs to be dumped and get loaded for later sessions’ usages.

Localization is more important in terms of world-locked rendering, while mapping is mainly used as supplement for localization. The map data needs to be sparse for efficient pose estimation, thus it cannot be used directly for downstream tasks. One of the common task is surface reconstruction, which builds meshes for the surface of the space, and find plane for physical space-aligned placement. Surface reconstruction requires dense map points, which comes from posed depth frame instead of the SLAM map. In this sense, SLAM is a upstream task and aims to provide pose while reconstruction needs to build the dense map itself and perform meshing.

For really large scale reconstruction, the task are usually performed offline, where SLAM may not be needed, and a similar technique called Structure from Motion could be used instead.

Metrics: Accuracy, Efficiency, Robustness

If a SLAM is needed, the next thing is to consider which 2 of 3 metrics are most important to your system since achieving the 3 metrics is difficult and may not necessary.

  • Accuracy + Efficiency: this means you could optimize your system for some specific environments, so system may degrade in other environments
  • Accuracy + Robustness: this means you could perform larger scale computation, so the output is not that realtime
  • Efficiency + Robustness: this means you performs very fundamental computations, so the accuracy is kind of low

Sensor Choices

According to the requirements and the target metrics, you could choose what types of the general sensors you want to use.

  • IMU: usually needed to provide motion info, but the price range is very large, thus the accuracy varies
  • Cameras: low lost but not very robust. You could also choose how many cameras you need
  • Lidar: high cost but accurate and could decrease the compute needs. You could also choose how many lidars you need

For specific product, you could also select other sensors. Some examples choices for different products are as follows.

  • Autonomous vehicle: IMU, Cameras, Lidars, GNSS, Speedometer, Radars
  • AR/VR devices: IMU, Cameras
  • Robot vacuum: IMU, Lidar, Camera
  • Drone: IMU, Cameras

Every sensor has their processing method.

  • IMU
    • Usually modeled as orientation, position, velocity, accelerometer bias, gyroscope bias
    • Use kinematic equations to propagate states
    • Could be fused with other info with ESKF (Error State Kalman Filter), or graph optimization with pre-integration techniques
  • Camera
    • Has different model imaging model to use. Usually is pinhole model plus different distortion correction
    • Usually needs feature extraction (with image pyramid) & feature matching to build data association
    • Optimize to minimize re-projection errors
    • Triangulation to get 3D landmarks
  • Lidar
    • Store the point clouds in k-d tree or voxel for nearest neighbor searching
    • Optimize to minimize point-to-point, point-to-line, point-to-plane errors

Optimization Choices

Tradeoff between accuracy vs efficiency.

What states to optimize

According to the requirements and the quality of the sensor, we could choose how many past states (pose and landmark position) we want to optimize. If choose not to optimize all states, the older states need to be marginalized.

The more states to be optimized, the more accurate of the system, but more computing.

Choices include

  • Only optimize current state: filter method (EKF, IEKF). Need to marginalize all past info
    • If also choose to marginalize all 3D landmarks, MSCKF needs to be used
  • Optimize current state and some past states: sliding window optimization method. Need to marginalize oldest info
  • Optimize current state and all past states: full optimization method. No marginalization.

How to choose linearization point

Linearization point is the value for the state which is used to compute Jacobian matrix. We could choose which states needs to use the updated value to recompute the Jacobian matrix, and which states could just use the FEJ (First Estimate Jacobian), which means only computes the Jacobian matrix at the initial linearization point.

The more relinearization, the more accurate of the system, but more computing.

  • For marginalized states, they use FEJ
  • For optimizing states, we could choose to fix their linearization point dynamically
    • Stop relinearization for some landmarks after some iterations since the gain to relinearize them is small
    • For states update: EKF vs IEKF

Implementation Tricks

Summary Description Examples
Base on practice To determine which result is the best, calculate their performance on some metrics and choose the best one 1. Choose H or F matrix during initialization
2. Use ransac to get results from inliers
Relative best Choosing best result needs both absolute and relative threshold 1. Determine loop closure
2. Relocalization
Coarse to fine Step 1 to calculate rough result quickly, then step 2 to refine it 1. 2-stage tracking
2. Add lots of mappoints and keyframes, then cull them
Image pyramid Extract features from different scale, and always match features in different scale ~

Map Management

Different-level map management may be needed based on different targeting scenarios.

  1. Data structure to store map points in memory
    1. Points could be associated with keyframes, and co-visibility graphs could be used to fetch semantic “nearby” points
    2. Voxel with hashing could be used to fetch physical “nearby” points efficiently
    3. Keyframes are also stored into the map, since it could group the points around each frame’s position
  2. Mappoints generation and culling
    1. Mappoints are generated by triangulation across different frames, and get inserted into map. Mappoints could have associated descriptor, confidence, and other fields
    2. Duplicate mappoints could be culled or merged into nearby existing mappoints. So that a physical position only has 1 matching mappoint
  3. Map update
    1. As new observations comes in, the position and properties of mappoints will be updated
    2. When loop closure happens, the whole mappoints may have dramatic update, while the value of pose needs to stay unchanged
  4. Local map retrieval
    1. For frontend tracking, it needs to retrieve the “local map” from backend. Instead of the geometrically nearby map (which could be achieved using hashing voxels), a local map containing useful elements for current frame’s tracking is more preferred, which could leverage both the geometrically and keyframe covisibility to fetch them
    2. After getting the local map, the map tracking will build the 3D points from local map <-> 2D features from current frame correspondence, which helps for limiting the frontend drifting. This process could be done by projecting the local map’s points into current frame.
  5. Map dump and reload
    1. After each session, map data needs to be serialized and stored into disk, and it could be loaded and get deserialized in the next session
    2. Each new session could build up a fresh map from scratch, and the reloaded map needs to align with current pose based on place recognition result. Therefore, the map doesn’t have absolute position and only stores relative position in terms of session origin

Others

  1. Active feature search for fast feature matching: active feature search means to search in an area on the image to get matched feature with previous mappoint. Compared to feature extraction and then brute-force match, active feature search is more efficient.
    1. Search in a region by grid
    2. Search with BoW tree
  2. Set bag flag for mappoint and keyframe
    1. When to set
    2. Always check the flag when using them
  3. Mappoint properties
    1. Distinct descriptor
    2. Normalized observation direction (could deal with occlusion)
  4. Mappoint fusion
    1. Maintains mappoints with voxel hashmap
    2. Fuse mappoints if 2 mappoints in the same voxel shares similar descriptors. Updating by merge their properties
    3. Leverage voxel to select mappoints candidate while doing the active feature search.

Reference

  1. 视觉惯性SLAM:理论与源码解析