← Research
metadome.ai — drawing-to-3D

Scoring the part, not the picture

24th September, 2026

Novel view synthesis generates an image of an object from an unobserved viewpoint, conditioned on a reference view, including the surfaces that reference never showed. Standard image-similarity scores cannot tell a correct mechanical part from a plausible wrong one. What we built to measure it instead, and what that measurement shows.

Takeaways
  • The generated views are enough to build the part. Three of five heavy-equipment parts and both brackets came out as recognisably the same component, with bracket thickness within 1–3% of the source CAD.
  • The usual way of scoring a generated view does not work. A render that filled a deep mounting socket with a solid block still scored 0.85 out of 1 for shape similarity. The part it describes could not be manufactured, and the score could not tell.
  • So we check feature by feature instead. Every hole, edge and corner in the correct answer is looked for individually and reported as found, missing or made up. Each run is also scored against deliberately broken models, so a number always arrives next to the score a failure would earn.
  • The model leads on two of the four ways of counting a feature, the edge and corner measures, by the widest margins recorded, and sits mid-field on the other two.

Every score is published beside the baseline a degenerate model scores on it. A number without its null is not a result.

1The problem

A model is given one view of a part and asked for another. Did it give back the same part?

The failures that matter are not ugly images. They are plausible ones: a bolt hole in the wrong place, a rib that vanishes between views, a deep mounting socket filled in as a solid block. Each produces a render that looks entirely reasonable and describes a component that would not function. This is a correctness constraint against a specific engineering artifact, not a fidelity constraint against a photograph.

Six panels: input view, ground-truth render, model output, aligned overlay, detected ground-truth features, and the per-feature verdict
One case, end to end. A part given from one oblique view and asked for another. The answer (b) has a deep hollow socket where the tool mounts; the model returned a solid block (c). The silhouette overlap after alignment is 0.85, a high score for a part that could not be manufactured. Panel (e) marks the eight features the answer contains, and (f) reports each one: two matched, six missing, seven invented.

Similarity scores cannot express that. Compressing an image into one number makes a missing bolt hole a rounding error: across eight standard shape metrics, a feature displaced by up to 12% of the image width produced no detectable change. The photometric metrics used in view-synthesis work fail differently. They reward not moving the camera, because a returned-unchanged view keeps plausible shading while a genuinely rotated one does not. In one run a total no-op outscored a real success on both PSNR and SSIM.

Two line charts: eight similarity metrics flat at zero across all displacements; correspondence methods rise sharply
Moving a feature and asking whether anything noticed. Take a correct render, displace one interior feature, and leave the outline untouched. (a) Eight standard similarity metrics register nothing at any displacement tested. (b) Methods that establish correspondence instead of comparing wholesale respond immediately.

2What we measure instead

Detect features in the reference and the candidate, match them by position after normalising framing, scale and in-plane rotation, and report each as matched, missing or invented. The output is a set of per-feature verdicts you can point at, not a score.

Three disciplines travel with it. Every run scores degenerate models, a flat outline and a fitted ellipse, that any usable metric must fail. Models are compared paired by case, because the part drives far more variance than the model does. And results are reported under several definitions of what counts as a feature, because that choice is not neutral.

One render shown five times: the part, then edge blobs, circles, enclosed regions and corner keypoints marked on it
Four definitions, one render. The same part yields 8, 4, 1 and 15 features depending on what you decide a feature is. The circles are visibly wrong: four overlapping rings that correspond to nothing on the part, a failure obvious in the image and invisible in a table of numbers. This is why results below are reported under all four rather than combined.

3Results

Six models, scored under each of the four definitions of what counts as a feature. The parts are real production components, none of them seen during training.

The definitions disagree. Feature recall under each of the four definitions, on one shared scale so the panels are comparable. Ours (highlighted) leads canny and corners; on circles and enclosed it sits fourth and third while gpt-image-2.5 leads. No model leads everywhere, which is why no aggregate score is reported. The enclosed panel is short for every model: true through-holes are the hardest of the four to recover.

4From views to a part

Feature recall is a proxy. What the pipeline has to deliver is the part itself in 3D, not pictures of it, so the question that settles this is whether the generated views are enough to build one. That is a different question from any of the scores above, and the one the product turns on.

They are, for most of what we tried. Three of five heavy-equipment parts and both brackets came out as recognisably the same component, with the mounting holes, slot patterns and bosses present and roughly in place. On the brackets, where the original CAD was available to measure against, the proportions matched the source to within 1–3% on the thickness axis. The two that did not were a part whose generated views were visibly degraded to begin with, and a near-planar cutting edge: a flat part seen edge-on carries almost no information to build from.

Where it works, the defects that remain are positional rather than structural: the right features, in approximately but not exactly the right places. That is the difference between a part someone recognises and a part someone can manufacture, and closing it is the current work.

EVERY MONTH ◈

One result, in public.

Model drops, findings, and hands-on demos, published monthly, no gatekeeping. Bring your hardest part.