
At 8 o’clock this morning, Apple’s iPhone 18 Pro series officially went on sale, and there are already long queues in many stores. Since the most eye-catching iPhone Duo has to wait until October, the popularity of the two models launched today, iPhone 18 Pro and iPhone 18 Pro Max, is relatively conservative. Especially for the former, the price on e-commerce platforms has already broken through during the pre-sale period.

Just three days ago, iOS 27 finally ushered in the official push, and then the rollover occurred immediately. A Reddit user handed a photo of a horse to Apple’s new Spatial Reframing feature, expecting it to be refinished, but instead a stranger showed up.

It is neither him nor the relatives and friends photographed in the album, it is just a strange “person”, a non-existent hallucination “person” from AI.

▲ Left: original image, right: Spatial Reframing modified image from Reddit user @friendofmany
It can be seen that in the original picture, there is a small piece of what looks like a hat on the back of the horse. When Spatial Reframing moves up and expands the picture, a man who has never appeared before is added.
It sounds so dangerous, how can the iPhone do such a thing – don’t be afraid, if you plan to buy the iPhone 18 Pro series, it brings the latest Apple Reference Image for sniper AI.
But wait a minute, is it possible that when iOS comes to market equipped with a large number of generative AI functions, Apple also launches a function that is the opposite of generative AI?
Yes, that’s what we find on the iPhone 18 Pro series, a seemingly contradictory set of new features.
A person grows out of a hat
Spatial Reframing is a new generative photo retouching feature in iOS 27. It has been available with the beta version since June. Users can not only crop photos, but also drag the image to simulate the viewing angle after the photographer moves sideways.
But after all, the camera doesn’t really move. Spatial Reframing wants the camera to change its position after shooting, so it must draw the area not covered by the original lens. Where there are no pixels, pixels are generated; if there is a person missing under the hat, a face and a body are added to it.
The principle of this imaging technology is still generative AI, so “making something out of nothing” is inevitable. TechRadar testers also found that when dealing with buildings and busy streets, the AI occasionally added windows and flanges to buildings that didn’t exist in reality. Compared with ordinary image enlargement, Spatial Reframing does not just fill in more background around the photo. It simulates the movement of the camera position, reprocesses the relative relationship between the foreground and the background, and then fills in the areas that were originally blocked or did not enter the lens after the perspective changed.
Apple can continue to improve perspective, light and shadow, and object continuity to make the results more and more natural, but it cannot eliminate the fundamental contradiction of this feature: as long as the camera does not record those pixels, the model can only guess what should be there based on the existing picture.
To be honest, mobile phone photography is not about light passing through the lens and falling into the photo album intact. At the moment you press the shutter, a series of links work together: HDR combines multiple exposures, Nightscape mode brightens details in the dark, and Portrait mode calculates the distance between the subject and the background. The photos we see are always the result of sensors and algorithms working together.
However, in the past, these calculations were mainly responsible for making what the camera has already captured clearer. Spatial Reframing crosses another line: it needs to fill in content for places that the camera does not see.
As a result, the photo still looks like a photo, but it no longer naturally equals a live record. The stranger did not leave the process of entering the frame, nor could he find the moment when the shutter was pressed; he directly transformed from a small piece of hat into a living “person”.
Apple leaves second testimony for photos
However, Apple is obviously not unprepared. The new iPhone 18 Pro series will be equipped with Apple Reference Image.
This is an optional Reference shooting mode. When enabled, the camera will leave a DNG “digital negative” directly signed by the sensor when shooting, recording the original pixels, sensor information and shooting time, and binding it to the main photo processed by conventional computational photography.
▲ Picture from: Instagram user @iJustine
This set of proofs is established when the phone leaves the factory. The main camera sensor generates a unique encrypted identity and binds it to the phone’s Secure Enclave. When the user presses the shutter, the sensor enters a special safe mode that directly signs the pixel data before the firmware modifies the screen; Private Cloud Compute then checks whether the sensor, phone, and timestamp match.
The system saves both a normal, editable photo and a secure digital negative signed by the sensor. The latter will not immediately become a viewable Reference Image; only when the user needs to verify it, it will be developed into an immutable reference image through Private Cloud Compute.
When it was first released, it was understood to be a “watermark” to identify the AI, but in fact, it provides an independent reference that can be used to compare what happened between the final photo and the sensor recording.
Take the horseback photo taken by Reddit at the beginning as an example (obviously he is not using the iPhone 18 Pro series), use the normal shooting mode, and then use Spatial Reframing to change the perspective, then what will appear in the album is just a complete photo that has been generatively processed. It may retain the editing history of the Apple album itself, and you can click “Restore”, but there is no sensor signature, and there is no Reference Image for external verification.
On the iPhone 18 Pro series, Reference mode does not immediately add three more photos to the album. When the shutter is pressed, two processing paths start at the same time: the conventional computational photography process generates a master photo that can continue to be edited; the sensor completes the signature of the captured pixels before the firmware modifies the picture, and saves it as a secure digital negative associated with the master photo.
This DNG negative is not yet a Reference Image. Only when the user requires verification is it fed into Private Cloud Compute, where demosaicing, tone mapping, and compression are completed in an auditable environment, resulting in a viewable reference image bearing Apple’s signature.
If the user then uses Spatial Reframing on the main photo, the album will show the edited version first. The main photo before editing is still retained in the non-destructive editing record, and the Reference Image is retained along another path and can be recalled through the Reference mark.
Thus, three levels emerged in the same photo project:
- The edited finished image shows a stranger standing behind the horse;
- Undo edited photos, no one;
- Reference Image generated from sensor records (needs to be called), no one.
The first two records how Apple processes photos, and only the last testimony can prove: the sensor never saw this person at the moment the shutter was pressed.
Is this unnecessary? No
Today’s technology, “original drawing” has become a concept that involves carving out a boat and seeking a sword.
If the Reference mode is not turned on when shooting, even if there is an “original image” in the album, all modifications and edits can be undone and restored in the Apple album, but there will never be an original record signed by the hardware that can be handed over to others for verification.
In front of the finished image produced by Spatial Reframing (even after modification by any generation tool), the “original image” is just material – and these materials can also be generated.
Reference Image pushes the starting point of trust to the sensor, so that the “original record” is no longer just claimed by the photo album, camera application or operating system.
After entering Reference mode, the sensor reboots and enters a dedicated safe shooting state. It converts the received light into pixel data and completes the cryptographic signature directly within the hardware before the data leaves the sensor. By the time the operating system receives these pixels, a shooting record has been formed that cannot be replaced without trace.
However, sensors are only the beginning of this chain of trust. Subsequent device identity, shooting time, and image development must be verified by Secure Enclave, encrypted timestamp, and Private Cloud Compute respectively. If any link cannot correspond to the original sensor signature, the final Reference Image cannot maintain a valid verification status.
Therefore, the Reference mode does not dwell on the concept of “original image”. It re-establishes a boundary with the sensor as the core of the fact, and builds a set of encryption and verification pipelines around this hard boundary.
Within the boundary, there is pixel data formed by light signals; outside the boundary, it may come from PS stitching, Spatial Reframing calculation supplement, or it may come from another set of AI generation tools.
Future cameras will deliver two photos and three states
In the future, photos may no longer only be divided into “original photos/retouched photos”, but will have three more detailed states:
- Without a Reference Image, the source of the shooting cannot be proven;
- A Reference Image exists, but the final photo has been edited;
- A valid Reference Image exists, and the final photo is essentially consistent with the sensor record.
In the past, when talking about how AI will change hardware, the answer was usually adding neural network engines, more memory, and better cooling to allow the device to run models locally.
Reference Image shows another impact: when software has the ability to reconstruct the results of hardware acquisition, the hardware needs to establish an additional boundary that is not dominated by software.
The main sensor of a mobile phone camera is not only responsible for sensing light, but must also have its own encrypted identity to prove that “the above pixels are from me” before any algorithm intervenes.
This is equivalent to letting the camera take on two sets of tasks at the same time: one process is responsible for interpreting, modifying and even generating the picture, delivering the best-looking and most usable photo; the other process starts from the sensor signature, only allows verifiable processing, and finally delivers a testimony that cannot be rewritten while retaining the certification status.
In order to achieve this, Apple does not simply add a watermark button to the photo album. The sensor of the iPhone 18 Pro series will generate its own encrypted identity during factory initialization, and then bind it to the secure area of the phone; after entering Reference mode, the sensor will directly sign the pixel data before the firmware starts to modify the image… In this set of encryption and verification processes, the relationship between the camera’s software and hardware changes.
Perhaps this dual-track communication will not only happen with cameras in the future – when AI can clone sounds, the recording device may also need to be able to leave the original sound track signed by the hardware; when the model can summarize and correct the data collected by the wearable device, the medical device may also need to keep an underlying record that the model cannot cover.
In other words, future hardware may have to deliver two results: one is the finished product that is sorted, modified or even rewritten by the algorithm for people, and the other is the negative that starts from the hardware acquisition and only undergoes a verifiable process. Although the Reference Image cannot be said to be an “absolute reality” that is completely uninterpreted by the machine, it also undergoes demosaicing, tone mapping, and compression; the real difference is that these processes can be inspected, but they cannot make up what the sensor did not see.