The fact that a significant portion of software today is written by machines is no longer news. In many development teams, code has long since ceased to be written line by line by hand. Anyone who sees this as the real game-changer has already missed the point.
The paradigm shift I want to write about here is a different one. And it recently took on a very concrete form for me: I saw a system whose output is no longer source code.
Three levels, and what’s happening right now
A brief overview for anyone who doesn’t work with software on a daily basis. For decades, software has been developed in three layers. At the top is a high-level language—Java, Python, C#—written for humans, readable, and open to discussion. At the very bottom is machine code: sequences of numbers that a processor executes directly. In between lies assembly language, which is essentially nothing more than a human-readable notation for machine code—a command usually corresponds to a machine instruction.
The compiler operates between the upper and lower layers. It translates what humans have written into what the machine executes. This division of labor is about seventy years old and has been the most stable constant in computer science during that time.
What I saw cannot be categorized within this framework. It wasn’t a high-level language. It wasn’t assembler. Nor was it what a standard processor accepts as an instruction. The processor and the model that generated it originated from the same development, and the user’s request was fed directly into the hardware. I can’t say what I was looking at—I can say that it lay below anything I’ve ever encountered in thirty years of working with software, and that at no point along the way did anything emerge that a human could have read.
I am aware that I am presenting this point as a personal observation and not as the current state of published research. Publicly available research takes it a step further: there, the processes generate assembly language—and assembly language, for all its cumbersomeness, is still a form of writing for humans. One command, one line, one readable abbreviation. As long as the result is assembly language, the chain remains tight, but it remains closed.
What I’ve seen doesn’t involve this stage. The result isn’t a written form, but rather what the electronics execute directly—without a human-readable version ever having been created at any point along the way. Not as an intermediate step, not as a byproduct—not at all. One might consider this an isolated case. I consider it a precursor.
Why the Compiler Was More Than Just a Tool
It’s worth pausing for a moment to ask what we’re actually losing. After all, the first counterargument is obvious: No one has ever read the machine code anyway. No one checks compiler output. We’ve been living with an opaque translation layer for decades—why should another opaque layer be a problem?
Because the compiler has two features that AI does not have.
It is deterministic. The same input produces the same output—today, tomorrow, and five years from now. And it has been rigorously tested over decades; there is even a formally verified version for a C compiler whose correctness has been mathematically proven.
Above all, however, the source code was preserved. We could restart the translation at any time and cross-check the results. That is precisely why the compiler’s opacity was tolerable—there was always an authoritative original from which everything could be derived.
If both of these elements are eliminated, the chain is no longer closed at any point. A stochastic, paid generator replaces a deterministic, free translator. And there is no longer an original against which to verify the translation.
What Dies, What Survives
This is where it gets practical for companies. It’s worth clearly distinguishing between what’s actually being eliminated and what’s staying—because the two are often lumped together.
What survives: Anything that tests behavior. A test that inputs a calculation and checks whether the correct result is produced will continue to work. Integration tests, end-to-end tests, business acceptance tests—all of these are independent of what the system looks like on the inside.
What’s going away: Everything related to the source code. Today, a developer can pause a program, go through it line by line, and observe how values change. If an error occurs, they receive a message with the filename and line number. Tools automatically check the code for typical error patterns before it even runs. Changes can be compared, reviewed, and assigned to specific individuals. None of this has any meaning if there is no source code.
That leaves exactly two points of contact: what I put in at the front and what comes out at the back. In between: nothing.
The point that gives me the most pause
The procedure itself is even more burdensome than the testing.
Today, the smallest unit of change is a single line. An incorrect tax rate, a missing condition, a rounding error—you change the relevant part, recompile, and you’re done. The change remains local. Everything else stays exactly as it was.
Without source code, the smallest unit of change is the system description. You modify the requirement and have the system regenerate. The result can turn out differently in any number of places, even where you didn’t intend to make any changes. No small change stays small.
For a system that has been in operation for seven years and requires a minor adjustment, this is a structural difference.
Four Objections—and Why They Don’t Convince Me
I’ve discussed these ideas with a few people. The same four counterarguments come up again and again, and they’re all valid. Nevertheless, I don’t consider them tenable.
“Then you just have a readable version generated as well.”
Technically possible, but economically unlikely. Anyone who takes this approach pays three times: for generating the machine code, for generating the human-readable version—and, on top of that, for the people who are actually able to understand that human-readable version. Yet it is precisely these personnel costs that are the reason this approach is being taken. A streamlining effort is being carried out, while at the same time continuing to fund what was supposed to be streamlined.
And even if you did: The readable version would not be the source, but rather a description. No one can guarantee that it matches the code that was actually executed. Anyone planning a change based on it is planning against a document that may differ. These are the costs of source code maintenance without its reliability.
“Regulated industries will prevent that.”
That’s true today. In medical technology, aviation, and automotive engineering, it must be possible to trace how a requirement became a function and how that function was tested. Machine code without source code does not meet this requirement.
But here’s the thing: These industries face the same cost pressures as everyone else. And traceability rules aren’t a law of nature—they’re agreements. They came about because they suited the technology of the time, and they’ll be adjusted as technology changes and economic pressure becomes strong enough. Anyone who believes that regulation can permanently hold back a development that reduces costs everywhere else is underestimating just how flexible regulation is.
“You can always extract readable code from machine code.”
That’s right, and AI models, of all things, are good at this. With the right tools, you can reconstruct something from a binary file that looks like source code and can be read.
And things are improving rapidly. The first open model, which was trained specifically for this purpose, appeared in 2024 and significantly outperformed general language models such as traditional tools. A 2025 paper by the same research group makes another noticeable leap forward, even outperforming current general models. Nevertheless, even for the best methods measured today, about one-third of the reconstructed programs fail the corresponding tests.
Und darauf kommt es hier nicht einmal an.
The difference is nonetheless crucial. What you get is a reconstruction, not an authoritative source. Two reconstructions of the same program may differ. You get a perspective—not the truth. That’s enough for understanding. But it’s not enough for accountability, verification, or liability.
“Machine code is expensive to generate, but the compiler is practically free.”
That is the strongest objection, and from a purely mathematical standpoint, it holds true today. Low-level code consists of a large number of small units with no condensed meaning; generating it requires significantly more effort from a model than generating a high-level language. The compiler, on the other hand, performs the translation at virtually no cost.
However, this objection reflects today’s economic reality. The costs of automated generation have been falling for years at a pace that quickly renders any snapshot obsolete. And this development is taking place—not because it is cheaper, but because it can achieve in specific areas what a rule-based compiler cannot. What begins in such niches rarely stays there.
The question that ultimately decides the outcome
I suspect that the success or failure of this endeavor won’t depend on the technology, but rather on an incentive problem that everyone is familiar with from their own workplace.
The savings are realized today, during development. The costs arise in three, five, or eight years, during maintenance, troubleshooting, and testing. These are generally handled under different budgets and almost always by different people. Wherever this pattern occurs, investment is reliably insufficient—documentation and test coverage have been prime examples of this for thirty years, and everyone knows their value.
So the readable version won’t be provided not because it would be impossible, but because the person who is saving money isn’t the one who receives the bill.
What I would advise companies to do
For companies that use software but do not develop it themselves, three things are relevant in light of all this.
First: The description of your system is your true asset. Not the code. If you have a precise record of what your system is supposed to do—including the rules, exceptions, and edge cases—then you remain capable of taking action, regardless of how it was implemented. If this knowledge resides only in the minds of a few people and in the code, it becomes costly. Incidentally, this issue is already relevant today, completely independent of AI.
Second: Verifiability must be built in from the start, not added later. If you can no longer look inside the system, you need a system that keeps a detailed log of what it has done and that can be compared against known results. This is an architectural decision—not one that can be made up for later.
Third: The staffing equation works out—but not as expected. Because behavior-based tests are here to stay and source-code-based tests are being phased out, the need for people who can read code is decreasing. It is increasing for people who can describe precisely, in technical terms, what a system is supposed to do, and who can assess whether a result is correct. These are different types of roles, and they are often closer to the business side than to IT.
Same question, different field
The same question arises outside of software development wherever systems begin to make decisions on their own: It is not the degree of autonomy that determines the benefit, but whether every single step remains traceable and, in case of doubt, requires approval. This is exactly what AnyAgent is designed for—the automation of process steps, in which it is defined for each step what the system performs on its own and what is submitted to a human. We’ve described our overall approach to this topic under AI Services.
