Back to list

Research

Robot–Camera Calibration by Adaptive Weighting of the Calibration Dataset

Paper: Adaptive Pair Weighting for Robust Hand-Eye Calibration

Researcher
Hwa-Rang Kim (WIM Inc.)
Date
August 19, 2026

In short — Conventional calibration treats every measurement in the dataset with the same weight (w = 1). This research performs calibration with a different weight assigned to each measurement, improving accuracy.

01 · Summary

Not every measurement should count the same

Systems that mount a camera on the robot wrist require hand-eye calibration, which determines the position and orientation of the camera relative to the robot. The answer is computed from a dataset of dozens of measurements collected while moving the robot through many poses, and existing methods treat every measurement with the same weight (w = 1) without exception. In practice, however, measurement quality varies, and when poor measurements enter the computation with the same weight as good ones, the whole result becomes inaccurate.

This research drops uniform weighting and computes the calibration with a different weight wi assigned to each measurement in the dataset. Low-quality measurements receive low weights and high-quality measurements receive high weights, so a few bad measurements cannot ruin the result.

The paper calls this method adaptive pair weighting (APW). For rotation, a lightweight Deep Sets network predicts a weight for each motion pair; for translation, iteratively reweighted least squares (IRLS) suppresses outliers analytically. The network is trained in a self-supervised manner without any ground-truth calibration. On simulated data with heterogeneous noise, APW reduces the rotation error by up to 85.4% and the translation error by up to 56.7%, and on real far-range data it reduces the rotation error by 59.5% and the translation error by 30.4%.

Core idea

Instead of fixing the weight assigned to each measurement in the calibration dataset to a uniform value (w = 1), the calibration result is computed with a different weight assigned to each measurement.

Bar charts comparing uniform and adaptive weighting. Top: all ten pairs, including noisy ones, have weight 1. Bottom: the noisy pairs are down-weighted
Conceptual comparison. (a) Closed-form AX = XB solvers give every pair the same weight (wi = 1), so noisy pairs (hatched, red) contribute as much as clean ones and degrade the estimate. (b) APW learns to down-weight the noisy pairs. (Paper Fig. 1)

02 · Background

Under uniform weights, bad pairs pull the estimate as hard as good ones

The standard formulation of hand-eye calibration is the matrix equation AiX = XBi. Ai is a relative robot motion and Bi is the relative camera motion observed at the same moment, and N such motion pairs (Ai, Bi) form the calibration dataset.

Hand-eye calibration geometry. The robot end-effector moves from pose i to pose i+1, and a camera attached through the transformation X observes a fixed calibration board
Hand-eye calibration geometry. When the robot moves from pose i to pose i+1, the relative robot motion Ai and the relative camera motion Bi are measured. The end-effector–camera transformation X is the same unknown at every pose, and because the two paths (via the end-effector / via the camera) must describe the same relation, AiX = XBi holds.

Widely used solvers (Tsai–Lenz, Park–Martin and others) all share the structure below, minimizing a weighted sum of per-pair residuals, with every weight fixed to 1.

X = argmin Σ wi ‖ei(X)‖²Conventional : wi = 1 (for all i)

Here ei(X) is the residual of the i-th pair, that is, the error that measures how far a candidate solution X fails to satisfy the i-th motion pair (Ai, Bi), and wi is the weight with which that pair enters the computation.

These solvers differ only in the definition of the residual ei, the parameterization of X, and the constraint placed on X (orthogonality through an SVD, a normal-equation solve, or unit norm through the smallest eigenvector); the weight wi enters all of them at the same place. The paper shows that five closed-form solvers — Park-Martin, Tsai-Lenz, Chou-Kamel/Horaud, Andreff and Daniilidis — all take this form.

In real deployments, data quality varies widely with capture distance, viewing angle, lighting and motion blur, and under uniform weights a contaminated pair pulls the result as hard as a clean one. Contaminated pairs are hard to identify in advance, and collecting more data does not help because the contaminated fraction stays the same.

For example, in large workspaces such as automotive or aerospace assembly, the camera-to-target distance can change by a factor of two within one sequence, and corner detection becomes less accurate as pixel density drops. In an eye-to-hand configuration, the camera sees the target at an oblique angle from some robot poses, and automated pose generation cannot reject a poor viewpoint. Uneven illumination, lens distortion near the image border and motion blur add further sources of error.

03 · Method

Keep the solver, change only the per-pair weights

The core of this research is setting wi differently for each pair in the equation above. Because the weights enter the common structure of existing solvers (a weighted sum of residuals) as they are, the method is applied by changing only the weights while keeping the existing solver unmodified.

The basic principle for setting the weights is the constraint that every pair must be explained by one and the same transformation X. Under this constraint, low-quality data reveal two kinds of deviation. First, the robot motion and the camera motion contradict each other within a single pair. For example, AiX = XBi requires the robot and camera rotation angles to be equal, so a pair whose two angles differ is self-contradictory. Second, a pair departs from the consensus of the whole dataset. Since all pairs are linked by the same X and must follow a common geometric relation, the further a pair departs from the relation formed by the majority, the lower its quality can be considered. The degree of deviation is scored and larger deviations receive lower weights; the scoring can be implemented in various ways (rule-based, statistical, learned models and so on).

Implementation in the paper

  • Rotation: A 12-dimensional feature comparing the rotation axes and angles of each motion pair is fed to a lightweight Deep Sets network (17,729 parameters) that predicts per-pair weights. The network is trained with a bootstrap consistency loss: the pairs are split into disjoint subsets and the network learns to make the solutions from the subsets agree, so no ground truth is needed and weight collapse is suppressed.
  • Translation: Because this sub-problem is linear in tX, iteratively reweighted least squares (IRLS) analytically down-weights outlier pairs with large residuals, with nothing to train.
  • Computational cost: Training finishes in about 107 seconds on a single CPU core, so the method runs on robot control computers without a dedicated GPU.
APW architecture. A weight prediction network (feature extraction, Deep Sets encoder, pooling and projection, decoder, softmax) and an adaptive weighting solver block (weighted solver, bootstrap loss, IRLS translation)
APW architecture. Per-pair features → weight prediction network → weighted closed-form solver (RX) → IRLS (tX). (Paper Fig. 3)

04 · Results

85.4% lower rotation error in simulation, 59.5% on a real robot

  • Suppressing the influence of low-quality data improves calibration accuracy.
  • Existing solvers, equipment and data-collection procedures stay as they are and only weights are added, so adoption cost is low.
  • There is no need for an operator to pick out bad data or to recollect data.

Experimental results

Simulation — Data were collected in the Gazebo simulator with a 6-DOF robot of known kinematics and a calibration board, so the ground-truth transformation X* is exactly known for quantitative evaluation. We used 300 poses (299 pairs) with per-pair heterogeneous noise (70% at base noise κ = 50, σ ≈ 8.1°, and 30% at 5× elevated noise) and 20 validation sets.

MethodRotation error RXTranslation error tXImproved sets
Uniform Park-Martin (+ LS)2.627°16.12 mm—
Strongest baseline (PM + P1×IRLS)1.339°7.27 mm17/20
PM + APW (proposed)0.385°6.98 mm20/20
Bar chart of rotation error for each of the 20 validation sets. APW is lower than uniform Park-Martin in every set
Per-set rotation error. APW consistently outperforms the uniform Park-Martin (Uniform PM) baseline on all 20 validation sets. (Paper Fig. 4)
  • The rotation error falls by 85.4% versus uniform Park-Martin and by 71.3% versus the strongest baseline (P1×IRLS), and the translation error falls by 56.7%.
  • With only 20 poses (19 pairs), the rotation error still falls from 8.891° to 2.54°, a 71.5% reduction.
  • Applied to five closed-form solvers with distinct algebraic formulations (Park-Martin, Tsai-Lenz, Chou-Kamel/Horaud, Andreff, Daniilidis), the rotation error falls by 84–96% in every case, confirming that the solver can be kept unmodified (solver-agnostic).
  • Under pose-level heterogeneity (17% bad poses, 44,850 all-pairs), the error is 0.137°, an 87.7% reduction versus uniform weighting.
  • A model trained on Gaussian noise reduces the rotation error by 80–84% under other noise distributions such as Student-t and Gaussian mixtures, without retraining.

Real robot experiment — A Neuromeka Indy7_v2 6-DOF collaborative robot with a wrist-mounted ZED X Mini stereo camera observed a 6×7 checkerboard (30 mm squares). The Park-Martin result from 26 close-range poses (d ≈ 0.5 m) served as pseudo ground truth, and 20 independent sets of 26 poses each were collected at far range (d ≈ 1.3 m), where lower pixel density degrades corner detection.

Photo of the experimental setup. A collaborative robot with a wrist camera faces a checkerboard on a tripod at distance d
Experimental setup. The camera-to-board distance d determines the observation quality. Corner detection is precise at close range, while at far range fewer pixels per square degrade detection accuracy. (Paper Fig. 5)
MethodRotation error RXTranslation error tXImproved sets
Uniform Park-Martin (+ LS)8.687°24.14 mm—
PM + APW (proposed)3.522°16.81 mm18/20

The rotation error fell by 59.5% and the translation error by 30.4%. Under the same conditions, the other closed-form solvers and robust estimation methods improved the rotation error by at most 16.1%. Because the reference is a pseudo ground truth, these results show consistency with a high-quality reference rather than proven absolute accuracy.

WIM technology and products related to this research

Explore our vision calibration technology, AI vision robot arm and research robot arm set.