Wing Commander III shipped in 1994 across four CD-ROMs, and a substantial fraction of that storage was Mark Hamill's face. Between the space combat sequences, the game cut to filmed scenes on a set, with lighting and blocking and actors delivering lines, and the sales pitch was explicit about the arrangement: the video was the reason to buy the thing.
That was the standard proposition of the period. Night Trap in 1992 was almost entirely footage. Phantasmagoria in 1995 was shot over months with a full production crew and composited over rendered backgrounds. The 7th Guest put actors into a house built in a computer. For roughly five years, from the arrival of cheap optical storage to the point at which real-time rendering became presentable, the most expensive games being made anywhere were the ones with the most film in them.
Then it stopped, almost completely, and the received explanation is that the acting was bad.
The acting was frequently bad. It is not the reason, and Tenebris Somnia, an Argentine horror game from Tobias Rusjan and Saibot Studios that cuts from Famicom-era pixel art to live-action sequences shot with a professional crew, is a good occasion to say what the reason actually was. This is an image-and-sound reading of a technique that failed once and is being attempted again under conditions that have inverted, and the inversion is the whole argument.
Start with what the cut actually does to the eye, because it is not a change of scene. It is a change of representational system.
A sprite is not a small picture of a thing. It is a symbol standing for a thing, and the viewer's visual system processes it as such: a handful of coloured squares that denote a person by convention, in the way a written word denotes without resembling. Ernst Gombrich's Art and Illusion, published in 1960, gave the standard account of what happens next. The beholder supplies the rest. A schematic image works by inviting completion, and the completion is performed by the viewer out of their own materials, which is why a cartoon face can carry an expression that no photograph could and why the reader of a comic panel does not experience the drawing as deficient.
Photographed footage does not invite completion. It presents a specific person, with a specific face, lit in a specific way, and the viewer's job is to perceive rather than to construct.
Cutting between those two systems within a single work therefore does something more violent than changing the picture. It takes the image away from the viewer. Everything the player had been building for the last twenty minutes, their own version of this protagonist assembled out of forty pixels and their own imagination, is replaced by somebody else's decision about who that person is, and the replacement is not negotiable, because a photograph is not a proposal.
That is a real effect and it is available for use. In 1994 it was not being used. It was being suffered.
Here is the difference, and it is entirely a matter of what the rendered portion of the game was understood to be.
In the FMV period, the low-resolution game world was a limitation everybody in the transaction understood as a limitation. The developers wished it looked better. The marketing apologised for it by pointing at the video. The player accepted it as the price of the medium. So when the game cut to footage, the cut carried an implicit statement: this is the real version, and the rest is what we could afford to render. The video was the ambition and the game was the apology, and each cut back to the rendered world was a small deflation.
That is the structural reason the technique died, and it explains why it died even in the cases where the production values were high. Wing Commander III's video was competently made. It still made the game around it look like a compromise, and as soon as real-time rendering could deliver a face at all, the compromise became unnecessary and the whole apparatus was abandoned inside about two years.
Now consider the situation of a pixel-art game made in 2026.
The pixels are not a limitation. Any team that can ship a game at all can ship one with photoreal materials and global illumination, because that capability is in every commercial engine and requires no special expertise. Choosing a Famicom palette in 2026 is a positive decision made against a freely available alternative, and it costs something: a large audience will decline the game on sight.
Which means the cut to live action no longer carries the old statement. It cannot say this is the real version, because the pixel version was not an approximation of anything. It says something else, and what it says is up to the work.
Masahiro Mori's 1970 essay is the other half of the mechanism, and it is worth using precisely rather than gesturing at, because the popular version of it has become almost useless.
Mori's claim was about a curve. As a figure becomes more humanlike, a viewer's sense of affinity rises, until a point shortly before full human likeness at which affinity collapses sharply into revulsion, before recovering at genuine human appearance. His examples were a toy robot, an industrial prosthetic hand, and a corpse. The structure that matters is that the danger zone is not at the abstract end. It is near the top, where the figure is almost right and the wrongness is therefore specific.
A Famicom sprite sits at the far left of that curve and is completely safe. Nobody is disturbed by eight-bit representation of anything, which is why the aesthetic has been so thoroughly domesticated as charming, and why pixel horror games have to work considerably harder than their contemporaries to frighten anyone.
A photographed actor sits at the far right, where affinity is high because the figure is a person. But the right-hand side of the curve is adjacent to the trough, and horror cinema has spent a century learning to push a real face into it: with make-up, with lighting from below, with frame rate, with a held shot slightly longer than comfortable.
A cut from one end of that curve to the other, inside a single work, in a single second, is a traversal no other technique produces. The viewer has been sitting safely at the abstract end, doing their own completion, and is deposited without transition into the neighbourhood of the trough.
So the technique has a genuine mechanism, and the sensible question is not whether it works but how much of it a game can sustain.
The answer is: not much, and this is the actual judgment the piece owes.
The effect described above is a contrast effect, and contrast is consumed by repetition. The first cut to live action in a game like this will be the most effective moment in it. The fourth will be a convention the player has learned. By the eighth the two registers have become a single register with two modes, the traversal is expected, and the game has spent its best asset on making a stylistic point.
The announced structure suggests the team knows this. The live-action sequences are described as arriving at certain key points rather than as an alternating structure, which is the correct dosage and also the difficult one, because a professional film crew and a cast are an enormous fixed cost for material that should appear five or six times.
That cost is worth naming, since it is the practical reason nobody does this. A small studio that decides to shoot live action has to become a film production for a period: crew, cast, location, lighting, an editor. The economics only work if the footage is central enough to justify the expense, and the aesthetics only work if it is peripheral enough to stay shocking. Those two pressures point in opposite directions, and the FMV era resolved them by making the footage central, which is precisely what broke it.
The visual collision is the one the marketing leads with, and the audio collision is arguably harsher.
A Famicom-era score is made of a handful of simultaneous voices, each a simple waveform, tuned to a fixed pitch grid, with percussion approximated by noise. It carries no room. There is no space in it, no reverb tail, no sense of a physical location where the sound is happening, because there was no way to render one and no expectation that anybody would. The player's ear accepts this completely and stops noticing within a minute.
Recorded production audio carries the room in every frame. A voice recorded on a set brings the walls of the set with it: the reflections, the floor, the distance from the microphone, the low hum of whatever was running. The moment that arrives after twenty minutes of square waves, the player is not merely hearing a different kind of music. They are hearing that the scene is somewhere, which the game up to that point has never claimed.
That is the sharper of the two ruptures, because hearing is faster than looking and considerably harder to disbelieve. A viewer can hold a photograph at arm's length as an image. A room tone arrives already located.
It is also the part most likely to be got wrong, since it is invisible on a trailer and expensive to fix. Production sound that has been over-processed, noise-reduced into cleanliness and levelled to match the chiptune, loses precisely the property that makes the cut land. The correct choice is to leave the room in.
The counter-reading is the obvious one and it should be stated at full strength.
The objection is that this is nostalgia with a theory attached. FMV is a failed format with a strong retro following, pixel art is a nostalgic style with a stronger one, and a game combining both is aimed squarely at people who find the 1990s pleasant to think about. The formal argument above could be constructed for almost any revived technique, and its persuasiveness is a function of how well it is written rather than of whether the technique is doing anything.
There is a good answer, and it is empirical rather than rhetorical. Live action returned to games about a decade ago without any pixel art attached to it, in a body of work that is unambiguously serious: Sam Barlow's Her Story in 2015, Telling Lies in 2019, and Immortality in 2022 are all built on footage of real performers, and none of them is nostalgic about the 1990s in the slightest. They use video because video does something specific, which in Barlow's case is that a recorded performance can be searched, re-watched, and re-interpreted in a way a scripted scene cannot.
So the revival of live action is established and is not a retro phenomenon. What Tenebris Somnia is proposing is narrower and newer: not video as a medium, but the collision of video with its opposite. Barlow's games are entirely footage. This one is a sprite game that occasionally is not.
That distinction is the whole claim, and it is testable in a way nostalgia is not. If the cuts land, they land because of the traversal, and a player who has never touched a Famicom will feel them.
There is one more asymmetry worth registering, and it favours the game.
Horror is the genre where the beholder's share is largest. A monster the player has assembled themselves out of insufficient information is always more frightening than one they have been shown, which is why Silent Hill hid its town in weather in 1999 and why every horror director since Val Lewton has understood that the shot before the reveal is the one doing the work. A pixel-art horror game is, in this sense, a machine for maximising the beholder's share: the player has been supplying the monster's real appearance for hours.
And then the game shows them somebody else's.
Retro horror games have a specific and widely acknowledged difficulty, which is that the aesthetic is comfortable. Two and a half decades of pixel-art revival have thoroughly domesticated the Famicom palette: it now signals warmth, craft, and a certain kind of independent production, and none of those are frightening. Developers working in this style spend most of their design effort fighting the connotations of their own presentation.
The usual solutions are known and mostly weak. Extreme content, which reads as juvenile more often than as disturbing. Sound design imported from a more modern register, which helps but produces an audible mismatch. Or the sudden shift to a higher-fidelity rendering for a single sequence, which is the technique under discussion here in a milder form and which several small horror games have used.
What live action offers is the strongest available version of that move. A photographed human being is the one image a pixel-art game cannot domesticate, because the domestication runs on abstraction and a photograph has none. It is, in a fairly literal sense, the only thing left that can still surprise an audience that has decided sprites are charming.
There is an economic argument for the technique as well, and it is the one most likely to determine whether other small teams follow.
A studio of a few people that wants a genuinely affecting cutscene has bad options. Three-dimensional character animation of the quality required to carry an emotional beat is among the most expensive things in the medium, needs specialists, and looks worse the closer it gets to convincing. Two-dimensional animation at sufficient length is comparably expensive. Both are months of work per minute of screen time.
A day of shooting with a small crew produces several minutes of footage of real people, whose faces already do everything a face needs to do, for a fraction of that. The economics genuinely favour it, and they favour it most at exactly the scale where the largest number of interesting games are now being made. This is not a new observation, and it is roughly why the technique appeared in the 1990s in the first place, when the constraint was storage rather than headcount.
What has changed is that a small studio can now shoot on equipment that costs less than a workstation, colour-grade on the same machine that runs the engine, and edit without renting anything. Saibot Studios is in Buenos Aires, which is a city with a substantial film industry and a deep pool of crew, and a game like this is far more feasible from there than from a place with an expensive production economy and no local talent to draw on. The technique is cheap where the film business already is.
The last frame before the cut is a sprite, sixteen pixels tall, that the player has by now spent several hours filling in with a face of their own construction. The first frame after it is a face that belongs to a specific person who was in a room somewhere in Argentina on a specific day, lit by somebody, doing what a director asked. Everything the player made is gone, and what has replaced it is not theirs, and it is looking directly at the camera.



















