Enzo Markarian

I wanted an achievement 0.4% of players have.

The practice tool that would have helped only ran on Windows and Linux. So the next project picked itself: port it to macOS. I did not write the code — an AI agent did, over seven tickets and an unattended overnight run. What I did was decide the order, define what counted as done, and refuse the answers that didn't hold up.

The setup

Devil Daggers is a first-person arena game with a very short list of people who are good at it. I wanted one of the harder achievements, and getting there means practising specific segments over and over rather than replaying the whole run.

There is a tool for exactly that — ddinfo-tools, an open-source app bundling a practice mode, spawn editor, replay editor and custom leaderboards. It reads the running game's memory to show live stats: your splits, your gem count, your homing daggers, as you play.

It runs on Windows. It runs on Linux. It does not run on macOS, and I use a Mac.

That is the whole origin. Not a work assignment, not a course project — a game I wanted to get good at, and a tool that wouldn't run.

What I was actually doing

This is the part worth being precise about, because it is the part that transfers.

I cannot write C# against the Mach kernel API. I did not learn to during this project. An agent wrote every line of the port, and it is genuinely good code — better than I would have produced if I had spent a month learning to write it.

What the project needed from me was different, and it turned out to be the part that decides whether an effort like this lands or wastes a weekend:

DecisionWhat I chose, and why
Order Risk first. The memory-reading work went in ticket one, not last — it was the piece most likely to be impossible, and finding that out on day one is cheap while finding it out on day six is not.
Evidence Capture the real error rather than guess at it. When the app died with no useful log, the next step was making the failure produce a message — not proposing fixes for a cause nobody had seen.
Safety The agent could push branches and open pull requests, but never merge to main. Anything that reached the main branch got read by a human first.
Done Each ticket carried conditions checkable by running something or reading the resulting change — not "works well". A separate reviewer with no stake in the work checked each one before it counted.

Where the work happened

Seven tickets, each in its own fresh session, each finished only when the checks passed and an independent review agreed the change did what the ticket asked.

What I decided risk-first order · what counts as done · capture the real error · main stays behind a human The agent writes the change one ticket, one fresh session, no memory of the previous one The gate checks it the build runs · the change is committed · a reviewer who never saw the work agrees Rejected? The reason is written down and replayed into the next attempt retry, knowing what already failed

Four attempts across the seven tickets were rejected. That number is the useful one — not because failures are good, but because a rejected attempt writes down why, and that record turned out to be worth more than the code.

Three things that went wrong, and what each one taught

01  The reviewer was confidently wrong

One ticket was rejected on a firm technical claim: the code called a system function that, according to the review, does not exist on macOS. That is the kind of verdict that ends an approach.

It was wrong. The symbol does exist — it is an undocumented compatibility export, confirmed by listing the symbols in the system library directly. The work had been correct and was rejected on a false premise.

What I took from it: a confident answer from a reviewer is still just an answer. This is the same lesson as my Wine project, where an agent told me a feature worked and I knew from playing that it didn't. Twice now, the moment that mattered was refusing to accept a claim that contradicted what could be checked directly.

The route was changed anyway — to the documented one, because relying on an undocumented export is a bet on Apple not removing it. Being right about the fact and still taking the safer path are not in tension.

02  Two failures weren't failures

Two of the four rejections had the same shape: the work was already finished and committed by an earlier attempt, so the retry correctly found nothing left to do, changed nothing — and was failed for changing nothing.

Both times the code was already right. The harness was measuring the wrong thing.

What I took from it: half the recorded failures in this project were the measuring apparatus, not the work. If I had read the number without reading the reasons, I would have concluded the tickets were badly written and rewritten perfectly good ones.

03  I documented a restriction that wasn't there

Reading another program's memory on macOS requires administrator rights — a real constraint with no way around it for a locally built app. I had it written down that "practice mode needs admin rights."

Then I ran it without them. Half of practice mode worked fine. Applying a practice spawn is just writing a file, exactly like the mod manager; only the live-stats half reads the running game. One phrase had collapsed two different things.

What I took from it: I found this by using the thing, not by reasoning about it. The correction matters because the user-facing message is the difference between "this feature is unavailable" and "restart with admin rights and it works" — and telling someone the wrong one wastes their time.

Where it ended up

It works. On a Mac, against a running game, the tool finds the game's live stats block by scanning its memory — 624 MB across 1,189 regions — and renders live run analysis correctly.

Windows and Linux find that data by reading a pointer at an address a server hands them. There is no such address published for macOS, so the Mac build has to search for it instead, and then prove that what it found is really the thing and not a copy of the same bytes sitting elsewhere in memory. That difference is most of why the port was more than a recompile.

The port runs on my fork, and both changes went to the original project: the port itself (21 files, in review with changes requested) and — separately — a small fix for a crash that turned out not to be ours at all. The maintainer closed that one in favor of his own fix for the same crash, which is the outcome splitting it out made possible.

The crash worth mentioning. During testing the app died at the end of a run. The obvious assumption was that the new Mac code had broken something. It hadn't: the cause was in shared interface code, introduced by the original project's own library upgrade months earlier, and it crashes identically on Windows and Linux. Anyone hovering a graph in that tool has been hitting it.

It went back as its own small pull request rather than buried inside a thousand-line port, so it can be reviewed in two minutes and helps people who will never run the Mac build. Splitting it also means the maintainer can take the crash fix and reject the port, which is a real possibility a bundled PR would have taken away from him.

What I put in the pull request

The port PR says, in its own verification section, that the app aborted during testing. It names the cause, says it is not the port's fault, and then flags a second possible cause I could not rule out.

That is not modesty. A maintainer who finds a crash himself after merging trusts nothing else in the submission; a maintainer who reads it in the PR gets to decide with the same information I had. The claim I could back — the block located after reading 624 MB across 1,189 regions, live stats rendering — is stated exactly and no further.

I had originally drafted a stronger line, saying replay reading and injection had been exercised against a live game. Checking the record showed that was not substantiated, so it came out. Writing a claim is easy; the discipline is checking it before someone else does.

What I'd say in an interview

I wanted to get good at a game, and the practice tool didn't run on my computer. So I directed the port. I didn't write it — I can't write kernel-level memory code and I didn't pretend to. What I did was decide that the riskiest piece went first so we'd fail cheaply if it was impossible, insist that a crash with no log meant capturing the real error rather than guessing, and keep anything reaching the main branch behind a human read.

The most useful thing I did was disbelieve a review. It rejected working code on the grounds that a system function didn't exist. It does. That is the second time on two projects that the thing worth doing was refusing a confident answer that contradicted what I could check.

Every sentence there is one I can be asked follow-up questions about, which is the only test of a claim on a resume that matters.

The part that outlived the project

The record of what went wrong here — every rejection, with its reason, written at the moment it happened — became the input to a separate piece of work: a system that feeds a finished project's failures back into the tooling that runs the next one.

That is written up separately →