Metric depth
Model-scaled depth from panorama inputs. Independent physical-scale accuracy is not established by these demonstrations.
VISTAlabs RESEARCH NOTES / SPATIAL AUTHENTICATION
We are an authentication team. Reliable 3D understanding is the enabling layer for a deeper question: is this device observing the same physical place that was enrolled before?
VISTAlabs is the research project spanning our private 3D data engine, VISTA-based capability research, metric reconstruction, spatial verification and model compression for devices. This note describes the team’s research history and direction, alongside retained public demonstrations. It separates outputs we can show today from experiments that still need an independently reviewable protocol. Private data-generation recipes and training details remain private.
01 / THE QUESTION
Our starting point was authentication. Passwords establish something you know. A device key establishes something you have. Palm, face and fingerprint recognition measure something you are. We are investigating somewhere physically real you are as a possible complementary factor.
Is this device observing the same physical place that was enrolled before?
A room is not a biological trait. We apply biometric-style principles—repeatability, distinctiveness, template matching, change tolerance, liveness and privacy—to physical environments. We call this spatial authentication, physical-place verification or environmental authentication.
The desired representation is a persistent spatial signature: evidence about stable structure that survives a new camera position or moved furniture. The security problem includes false accepts, false rejects and adversarial capture. A recognizable room image alone cannot establish trusted presence.
02 / THE DATA WALL
To study the same place across sessions, we needed controlled variation. One beautiful capture tells us little about whether a room remains identifiable after its lighting, contents or observing device changes.
Useful supervision extends beyond RGB: metric depth, physical scale, camera transforms, surface and point geometry, normals, visibility, cross-view correspondence and scene topology. Those signals have to agree in a common coordinate system.
Real capture becomes difficult at the level of coverage. Scanner cost is one part; rooms, building layouts, architectural variation, repeat scans, camera trajectories, lighting, object configurations and device conditions multiply the collection effort. Maintaining geometry and calibration across that variation became our data bottleneck.
Real LiDAR is valuable for measured metric geometry, subject to sensor and calibration limits. Our team could not practically collect the breadth of LiDAR-grounded variation this research called for. We therefore built controlled simulation to supply similarly structured, physically scaled geometric targets.
“LiDAR-style metric supervision” describes the structure of those targets. It does not turn simulated points into physical LiDAR captures or establish equivalent accuracy. No real LiDAR capture is presented as evidence in these notes.
03 / THE DATA ENGINE
We built Blender- and Unreal-based tooling to create controlled indoor environments and generate training and evaluation data. The purpose was known geometry with deliberate variation: experiments in which we could change an observation while keeping the underlying place fixed.
Room dimensions, topology, ceilings, openings, object placement and materials. Vary lighting, clutter and occlusion without losing the scene’s geometric reference.
Camera position, height, orientation and field of view. Vary capture density and overlap while retaining camera transforms and physical scale.



Actual public demonstration outputs, illustrating supervision types. These are not private synthetic training examples. RGB, depth and normals share a cropped panorama grid; points use a perspective view. Displayed depth is a model estimate, not simulation ground truth.
The engine supports a verification-oriented experiment design. Positive pairs keep the physical room fixed while varying cameras, lighting, furniture or visibility. Hard negative pairs use different rooms with deceptively similar dimensions, furniture, materials or layouts.
Preserve architecture.
Vary the nuisance conditions.
Change physical structure.
Keep visual shortcuts tempting.
The intended lesson is persistent geometry, not “same couch” or “same wall color.” Known simulation geometry makes these questions controllable. It does not remove the synthetic-to-real gap; held-out real recaptures remain necessary.
04 / PRE-ASTRA
The project did not begin with a single model or a single launch sprint. Earlier assistants helped our team develop the infrastructure: Blender automation, Unreal tooling, procedural scene generation, camera placement and batch rendering.
The work also included metadata generation, depth and geometry export, dataset organization, validation and training preparation. These are the connections that let a scene become a usable, traceable training example.
Scene variation, cameras, render batches.
Observations, depth, geometry, metadata.
Coordinates, scale, organization, training inputs.
Team-reported project history. Proprietary generation code, corpus size and training records are not released in this note.
05 / ASTRA
As the system grew, failures became harder to localize. An error could originate in the model, a transform, scale, camera sampling, visibility, loss design, evaluation protocol or the gap between simulated and real captures. Useful assistance had to connect these layers.
Astra helped us reason about missing failure cases and weak camera distributions; separate appearance from geometry; design scale checks; propose targeted synthetic examples; compare experiments; organize evaluation; and refine the generation pipeline. A dataset problem calls for a different intervention from a model-capacity problem.
A convincing rendered room can hide sparse or missing geometry.
Inspect native points, camera transforms and the rendering stage separately. Compare the same camera before attributing the problem to the model.
Recorded launch work: matched cameras made representation changes interpretable; unseen coverage remained limited.
Failure in one view can reflect weak camera sampling, occlusion or a missing scene configuration.
Form a hypothesis, generate targeted camera and visibility conditions, then evaluate on held-out scenes. Change one explanatory factor at a time.
Research approach: targeted synthetic examples can test a suspected distribution gap. No new training gain is claimed here.
Higher input resolution can improve visible detail without improving every metric.
Hold camera and renderer settings fixed. Compare geometry, appearance and failure cases rather than one attractive screenshot.
Recorded launch work: resolution comparisons were mixed; detail selection was kept separate from a universal accuracy claim.
A model-derived distance can look plausible without an independent reference.
Define endpoints, retain scale and transform provenance, and compare against known geometry or an independently measured baseline.
Recorded launch work: measurements remain explicitly model-derived. Independent centimetre-level accuracy has not been established.
The overall workflow is the team’s account. Recorded launch checks support the camera, resolution and scale examples above; this page does not publish a controlled attribution study or new training benchmark.
06 / THE CAPABILITY MODEL
The first question was whether the required physical signal could be reconstructed reliably enough at all. A high-capacity VISTA-based research system gave us a reference point for studying metric geometry with our data and evaluation stack.
That reference model allows us to inspect depth, cameras and shared points before deciding which parts can be compressed. Downstream Gaussian reconstruction is a useful scene visualization stage; it is not the central authentication contribution or a substitute for geometric evaluation.

Saved VISTA geometry and cameras from one capture. 109,760 native points; the orbit is real camera motion around this representation. It does not demonstrate a second session, complete unseen surfaces or verified physical scale.
The large model establishes a capability reference. The dataset gives us a route to compression.
This is a research progression, not a claim that every stage is solved. Cross-session reliability, a verified matcher and consumer performance require their own evaluations.
07 / RESULTS & SCOPE
Our first goal is to determine whether ordinary visual captures can recover an environment consistently enough for verification. This release demonstrates the enabling representations; authentication performance is a separate evaluation.
Model-scaled depth from panorama inputs. Independent physical-scale accuracy is not established by these demonstrations.
Retained camera transforms and virtual source-view frustums. They are not an independently surveyed camera baseline.
Organized points and matched saved cameras connect the displayed representations of each scene.
A research objective across recaptures, devices and viewpoints. This release does not supply a validated multi-session benchmark.
The key requirement for persistent spatial signatures. Re-rendering one capture is not evidence of repeatability across captures.
The intended downstream task. No released matcher, authentication score or security decision runs in these demonstrations.
No centimetre-level spatial error or verification percentage is published here. A spatial error needs a defined metric, independent reference, sample count and evaluation split. Verification needs false-accept and false-reject behavior at a specified threshold, not a vague accuracy percentage.
Three accepted indoor demonstrations cover model-scaled depth, camera geometry and points. The public media are not a multi-session benchmark. Normals are derived from organized geometry. The model-distance annotation is not an error bound, and visual acceptance is not metric ground truth.
Unseen and polar coverage is incomplete. Sparse geometry, holes and view-dependent details remain. The 180° point orbit follows actual saved cameras; Gaussian novel views use a more restricted path. Matched representation comparisons use the same saved camera.
The single-photograph showcase on the main page is a separate RGB-only downstream reconstruction branch with its own depth and points. It does not consume VISTA depth. Final Gaussian renders use 1.5 million primitives; center displays use a 240,000-point subset. Primitive count is not a quality metric.
The 1.12–3.72 model m depth legend is a visualization range. Virtual source-view frustums do not establish a measured physical camera baseline; measurement overlays do not claim depth occlusion.
08 / NOT JUST 3D
A visually convincing reconstruction answers an appearance question. Physical-place verification must answer whether stable geometry can be recovered and distinguished across time, nuisance changes and attacks.
Does the same place align across sessions, devices and partial views?
Can architecture be separated from temporary objects and furniture movement?
Can physically different rooms be rejected even when they look similar?
What are the false accepts and rejects? Can templates remain private and replays be detected?
WindowOpeningCeiling structureWalls, room dimensions, corners, doors, windows, permanent openings, ceiling structure and built-in architectural features.
Candidates for a persistent spatial signature. Long-term repeatability must still be measured, including renovations and damage.
Concept illustration with explanatory annotations. It is not a segmentation output, training example or measured feature weighting.
Persistence is contextual. A wall may be remodeled; a large appliance may stay for years. We need to learn and evaluate which structure carries identity over time, rather than assigning permanent trust to an object category.
09 / SPATIAL VERIFICATION
The proposed flow turns camera observations into a compact spatial template, then compares a later reconstruction against persistent geometry. Privacy, freshness and device trust must be designed alongside the representation.
New viewpoint, capture position or lighting
Persistent architectural geometry
Repeat captures of an enrolled place, separated by session and device. Evaluate genuine comparisons at a fixed threshold.
Test-design scenarios, not live comparisons. No captures, template, similarity score or authentication decision are produced here.
A same-place match would still not, by itself, prove that an authorized person or trusted device is present. A screen replay, rendered scene or manipulated capture may reproduce visual evidence. Liveness, spoof resistance and trusted capture remain open parts of the research.
10 / THE CONSUMER PATH
A large model is useful for establishing a reference. The product path is a specialized system that preserves the signal needed for verification, rather than reproducing every capability of the research model.
Our private data-generation and evaluation pipeline can support distillation, smaller architectures, quantization, fewer input views, lower resolution and mobile optimization. Each reduction must be checked against geometry, repeatability and the operating error rates.
The intended consumer experience should not require a specialized LiDAR scanner. Local inference and compact private templates are design goals; mobile latency, memory, power and template protection have not been established by this release.
11 / BIOMETRICS
Our work in palm recognition and authentication led us here. Biometric systems already force us to reason about repeatability, distinctiveness, templates, liveness and privacy. Those principles prompted the question of whether a physical environment could supply a complementary signal.
Proposed multi-factor composition. Higher confidence would depend on measured errors, independence between factors and a defined threat model.
Spatial verification is not intended to replace palm, face or fingerprint recognition. It asks a different question about the observed environment. A useful system must be explicit about when that signal helps, when it fails and what it reveals.
12 / WHAT COMES NEXT
The data engine and geometry stack make the authentication question practical to study. The next work is to turn the idea into a measured verification system, then reduce its cost without losing the signal.
Collect independent recaptures across sessions, devices, lighting and object changes. Define alignment, physical reference measurements, visibility masks and error aggregation. Separate scene identities between development and evaluation, and report coverage alongside error.
False-accept rate (FAR) measures accepted different-place comparisons; false-reject rate (FRR) measures rejected same-place comparisons. Report both at a threshold fixed on development data. Include sample counts, hard negatives, confidence intervals and a receiver operating characteristic (ROC) curve. Equal-error rate (EER) summarizes an operating point where FAR and FRR coincide; it does not choose a deployment threshold.
Test how much a spatial template reveals, how it can be protected or revoked, and whether local processing is sufficient. Define attacks using photographs, screens, rendered rooms and manipulated capture streams. Evaluate freshness, trusted capture and recovery when a place changes.
Measure latency, peak memory, power, device variation and error-rate changes after distillation or quantization. Fewer views and lower resolution must be evaluated under the same held-out verification protocol.
The data was the bottleneck. The generation system made controlled experiments possible. Earlier assistants helped build the infrastructure; Astra helped structure the research loop; VISTA provides the high-capacity reference. Our direction is smaller, faster, private consumer models, with verification claims grounded in a reproducible protocol.
Follow the research ↗Explore the geometry demos ↗The Lythwood, Cayley and Artist Workshop public demonstration photography and panoramas come from retained Poly Haven CC0 source records. Reconstruction assets shown here are actual retained pipeline outputs. The architectural cutaway is a concept illustration, not a model output or sample from the private dataset.
The data-engine and pre-Astra history follow the team’s account. Private corpus size, generation recipes and training receipts are not published. The Astra examples distinguish the wider research approach from checks retained in the launch workspace. This note is not a peer-reviewed paper, independent audit or security certification.
Upstream software and model licenses are separate from the public demonstration imagery. No weights or private training corpus are distributed by this website.