Every augmented reality demonstration I have seen this year has the same object in it, and once you notice it you cannot stop noticing it. There is a printed square of black and white pattern on the table, and the system is looking at that.

Take the marker away and the demo ends. Not degrades. Ends.

This matters because the marker is doing the entire hard part while the presentation is about something else. The impressive bit appears to be the model floating in the scene, correctly sized, turning as you move the camera. The actual achievement is knowing where the camera is in space, and the researchers solved that by placing an object in the world whose exact dimensions and pattern the software already knows. Everything after that is arithmetic.

Remove the known object and you are asking a device to work out its own position and orientation from an arbitrary view of an arbitrary place, in real time, from a camera. That is not a harder version of the same problem. It is a different problem, and it is not close to solved.

Which is why I get uneasy watching this get discussed as though it is nearly here.

Look at what a current phone would need to do it. A camera good enough to track features in ordinary light, which today’s are not, particularly indoors. A compass, which almost nothing ships with. An accelerometer accurate enough to know which way the device is tilted, and most of the ones fitted are there to rotate the screen. Location, which is finally getting good enough to be useful, though a few hundred metres is not the precision this wants. And enough processing headroom to do computer vision on a live video stream without the battery going flat in twenty minutes.

Not one phone on sale has all of that. Several have two or three.

My guess is that the first version that reaches ordinary people will be much less ambitious than the demos and much more useful. It will not track surfaces or place objects convincingly in a room. It will use the compass and the location fix to work out roughly which way you are pointing, and draw labels over the camera view. That restaurant. That station. That is north.

Crude, occasionally wrong, and genuinely handy in a city you do not know. It needs no computer vision at all, which is exactly why it will arrive first.

The version in the research videos needs a phone that does not exist yet. Give it a compass, a decent camera and a few more years of silicon, and the conversation becomes real. Until then, when someone shows you this, look at the table.