An Interview with Doc Midnight
Preamble: Welcome to a new feature of the blog where I interview leading engineers and talk shop about their AI coding experiences and workflows 👾
If you’re an engineer and would like to be involved please reach out!
Doc Midnight is open to short-term contract engagements. You can reach out to me if you’d like an introduction.
Introduction
Tom: Thanks for jumping on. Before we kick off, what's the story behind the name Doc Midnight?
Doc Midnight: People started calling me the Midnight Doctor because I was always up at night fixing shit, and then they changed my global admin name to Midnight Doctor. The problem was, with how our old naming convention worked, it ended up being D Midnight. So, Doctor Midnight. That's how it came to be.
Tom: Let’s get started, how do you earn a living?
Doc Midnight: I earn a living two ways. Firstly, I work for a large financial institution, where I'm basically an FDE, a forward deployed engineer slash full-stack engineer. I build anything the business wants to build. I'm a technical person embedded within non-technical people. My living is earned by being able to turn vague business ideas into technology solutions.
Outside of my full-time job, I earn a living by taking on engineering work: building anything from AI agents through to complex data pipelines, whatever's coming through the pipe. Doing pen tests, doing AI red teaming, a lot of things. That's what I do.
Tom: Awesome. And then, what's your stack?
Doc Midnight: From a cloud provider perspective, I'm very Azure-focused. I have been, because that's where I started my career. I use AWS and GCP when Azure can't cover it.
I'm also a very big fan of on-prem. I've got a number of old-school Dell PowerEdges and a number of Lenovos in a rack under my house, along with some AI-specific hardware I've cobbled together over the years. A couple of Mac Studios, a couple of DGX Sparks, all sorts of random things like that. All of my stuff has a specific purpose. I own a couple of computers and a couple of Macs as well for my day-to-day.
From a software perspective, given the advancement of AI, I write most of my own stuff. Other than the standard stuff like Discord, which obviously I use out of the box, everything else I've written myself over the years for my own purposes. I'm a very big fan of open-source software that you can change and chop to your own usage. So yeah, that's my stack.
Tom: What's an example of something you wrote yourself?
Doc Midnight: A good example is that I don't like the out-of-the-box coding harnesses from a customisability perspective. So I took a fork of OpenCode and several others, took pieces from each, and wrote my own called Orro. Orro is basically what I want as a product, as a developer, as an engineer, and also as a cybersecurity professional.
Another example: I don't like Adobe generally, and especially not their PDF software. So I wrote my own PDF tool for all the functions I need, like PDF development and signing. I forked a couple of open-source projects and hooked them together.
Tom: I'll have to hit you up for that repo.
The harness
Tom: Moving on to the harness questions. What's your AI coding workflow? How do you approach it?
Doc Midnight: The same way I did before AI. I was a very big fan of spec-driven development. Not so much test-driven development, because I thought spec worked more with how I engineer.
I get up at about 7am and spend the morning writing all of my packets for the day. I do what's called a phase breakdown: I write down everything I want done in a big blob, and then I bust it down into phases. Then I run it through a pipeline that splits it up into agentic-friendly loops and tasks, so it can optimise for sub-agents in my harness. That comes back with a packet for me to review. Usually by 8 o'clock, I've got about 10 hours of work for the whole local stack to run through.
By the time I'm done with my BAU workday and I'm sitting down having my dinner, I'll review all the PRs for that day. I'll put changes in, ask for different things, and feed it back. For the most part, I just approve them, the ones that are just features and look good and look like they're working. There are obviously multiple steps all through that. But my overall workflow can be defined as: I set the requirements, and the agents do all of the implementation, testing, security work, all the stuff I used to do manually, end to end. Documentation uplift, the whole show, all at different stages, and then it feeds back to me for review.
At the end of each week, I take the traces from each session and feed them through a customised pipeline that helps enhance the harness. Be it skills, fine-tuning the models, adding tools, whatever needs to be done to speed things up and lower token usage. Over this year, most of the models I plug in, GLM, DeepSeek, whatever they may be, even non-tuned models, score much higher on benchmarks out of the box because of the way I've built this harness. That's just the nature of it. You can get three or four points on each benchmark, which is pretty cool.
Tom: That's a good point, because we've talked a bit before about how you evaluate the changes you make to the harness. How do you set up that evaluation framework?
Doc Midnight: I'm a very big Python guy. There are plenty of packages and plenty of different ways to manage it, but if you can't measure it, then it's not real. I also measure it with my own eyeballs, because the best AI assessor is your eyes.
I have a set of tasks that I run on every model. I also use Git pretty heavily to rerun a previous run and see if it's better with the changes. So there's a continuous improvement loop. Say I'm doing a coding workflow, or even a pen test on a dedicated box. If I've made a bunch of changes, I run the same workflow again, using Git to back my changes off so I can try again with no memory of the previous run, just the new tools. If it implements faster, better, more efficiently, across all the different metrics, then the change is worthwhile and I keep it. If not, I toss it out, or optimise it again and rerun.
I spend about a day a week on optimisation, but it's also allowed me to punch out three times as much code, which is pretty decent.
The other way is traditional benchmarks. Traditional data science benchmarks don't really help me, but things like the cybersecurity benchmarks and a bunch of the code quality benchmarks do. Still, I'm a very strong believer that human eyes are better than all of them, and your vibes are better than all of them. You can measure everything, but the models are so heavily optimised for the out-of-the-box benchmarks that you have to write your own for your own use cases.
Sign up for the next AI engineer interview
Once a month you'll join us and get deep into the weeds of the latest AI engineering with real practitioners.
Get the next interviewTom: Yeah, for sure. I think we could talk for a long time on that topic, so moving on. What's the best AI skill you've ever used or implemented?
Doc Midnight: If we look at "best", there are two: the most useful, and the coolest.
My all-time favourite, the one I enjoy the most, was recently shared with me, and it lets you create video games. I've been hammering that skill and wasting so much time on it. It's stupid the amount of tokens and compute I've burnt on it. I can't name it off the top of my head, but it's dream something? (Note from Tom - https://github.com/achimala/dream-loop) from memory, and it's absolutely fantastic.
The most practical one is a skill I used before I built my optimisation pipeline, and I still use it in my BAU job because we haven't got my harness at work. It's called Threadripper, or for people who don't use Amp, you'd call it a session ripper. Basically, it looks at all of your previous sessions, defines better skills and MCPs you can use to optimise your flow, and then tests them.
It was about a four-page skill, and it was heavy on tokens. I ran it once a week, same as my current optimisation, and it would spin up maybe 50 skills each time. Then I'd clean them up, and I had a global repository of skills that I keep going back to. I absolutely loved it. I still use it at work almost daily, because I find things every time I run it.
There's a good point there about skills: make sure you have a Git repo of all your skills and just pull them into every project. Have a one-shot command that runs at the beginning of each session that pulls them in, or at least updates them. That goes for all your documents, diagrams, all the stuff you personalise yourself. That's a really good tip for the younger player.
Tom: I haven't heard of Threadripper. I'm going to give it a red-hot go. On the topic of skills, how many do you have loaded at any given time? I think it's a bit of an untalked-about topic, whether there are too many skills or too few. I have a lot, actually.
Doc Midnight: I have one skill loaded, because all of my skills sit behind an MCP server. The MCP server, for lack of a better word, is basically my library. It's got all my skills, my scripts, etc. hidden behind it. So it's only one endpoint, and it can request from the MCP server and pull things in as needed. When the model hasn't got an answer or a talent, it knows there's one skill to look at, and that stops the token burn.
There are other ways to do exactly what I've just described. That's just how I do it, because it's how I like it. It's written in FastMCP, and it's a really nice little setup that's easy to deploy in most environments.
I've been spoilt by frontier models at work, because they've got such long context windows, so I just load everything in if I'm being lazy. But best practice is obviously not to do that. When you start using local models, you don't have God's grace to just do anything. Even with a 1 million token context window on local models, that's more of a pie-in-the-sky type vibe.
Tom: Yeah, just because you can doesn't mean you should. So what used to work great in your workflow, whether it was tools, skills, or how you did things, that doesn't work that well anymore, and you actually ripped out?
Doc Midnight: Look, I used to rely heavily on out-of-the-box memory management MCPs, things like Context7. I thought they were going to be the bee's knees, because a lot of what I do is around development or pen testing. A lot of those memory management MCPs are great, and I'm not shitting on anyone's products by any means. But when you start working with these tools, you find gaps that are specific to you. So you build your own.
I have my own memory management stack now, and I've pretty much replaced all of my external MCPs with ones I've written myself, because it's just better for me and for my harness. It's not scalable to other people, but for my own harness, I've got stuff set up specifically for it, because you need customisability. I'm a very big fan of customisability.
Tom: Do you think that approach makes sense for most people, given how much easier it is now?
Doc Midnight: It depends on the people. I'm a freak of nature. I love problems. I really, really love problems. My wife says that's probably the only reason we're together: her ability to generate problems. God love her. That was a joke meant to be in our wedding vows.
All jokes aside, the barrier to entry for that stuff is very low, but the drama with MCPs is that they're just a JSON wrapper around an API. If you're not really careful, you can make an absolute ass of yourself. We were doing a pen test the other week on a Copilot Studio agent that had a custom MCP, and you could bust out a hell of a lot of data, because they'd written their own MCP but hadn't set up the auth correctly. It was just read-write all. You could do some wild stuff on their DB.
Tom: If you could only pick two out of three, which would you pick: cheap, fast, or autonomous?
Doc Midnight: For my own stuff, it's cheap and autonomous. Speed means very little to me, because I'd rather things be done correctly than quickly.
In my current organisation, it's probably fast and autonomous. Cheap is less of a priority there, because the ROI is still extremely high. It comes down to where your ROI is. Everything comes back to return on investment for me. My stuff needs to be 100% before I'm comfortable with it. Some workplaces aren't following that anymore. I'm one of the few people left who likes good-quality code, I swear.
Rapid fire
Tom: Cool. The next five are meant to be one-liners, rapid fire. Terminal or IDE?
Doc Midnight: Over the years, I've become more of a terminal developer, purely because a lot of what I do is in headless environments, especially GPU clusters and things like that. Doing everything any other way is a pain in the ass. You have to use the terminal, or it's just as rough as guts.
Tom: Read every line of code, or ship if it runs?
Doc Midnight: Probably a medium between the two. I like "ship it if it runs", but I need to be able to understand the code that was written. Generally speaking, if I can understand what's written and how it all fits together at a high level, I let it go, as long as all the tests pass. That's a big one.
Tom: Voice prompting or typing?
Doc Midnight: That's an interesting one, actually. For the morning packets I mentioned, I walk around my office in a circle. I've actually damaged the carpet, so now I've got this little rug that I walk on. I have my coffee and I just talk, and then I clean up my notes and craft them. So I actually do like voice prompting. Typing is my main medium after that, because I'm out and about so much, but I really like interacting with voice.
I also like talking to these models while I drive, going back and forth. My current iteration of the harness can reach out and make a phone call to me if it's got a question. So on average, three times a day, I get a call through one of my Linux servers saying, "Hey, I've got this problem." It's just a nice voice model, and we go back and forth. So if you ever see me talking on the phone and you can't work out who I'm talking to, it's likely one of the models on my stack.
Tom: Has AI made you better or lazier?
Doc Midnight: That's a loaded question in a way. If I think of the sheer quantity of stuff I've been able to build, probably better. If I think about whether I'm quicker to reach for it than I should be, probably, yeah. So maybe a little bit lazier.
But I think it's the same as having a team. Before my current gig, I had multiple people working under me, and you become more of a leader than a doer. That was an interesting shift, and I treat AI like that. So yeah, I'm doing less. But I actively try not to use AI for anything other than development or security work. I use it for busy work, not for things that require my critical thinking.
Sorry, that was a long answer.
Tom: No, no, it's totally a loaded question. That was a good explanation. One mega-prompt or lots of small ones?
Doc Midnight: Depends on the task. I burn a full half a million tokens of context on my morning packet, because it's so detailed. Then it gets broken up into goals, tasks, and phases. That's how it cuts it up to make it acceptable for my agentic flow. So I'm a very big fan of War and Peace in the initial prompt, then breaking it down into very granular tasks. That's where the advancement in sub-agents has become so handy.
Back when we had to use AutoGen and similar frameworks, a lot of them didn't get the performance we now get out of the box from consumer tools. I'm not shitting on AutoGen or anything like that. I'm just saying the task breakdown and granularity is fantastic now. So one mega-prompt, then break it down into smaller tasks.
Closing thoughts
Tom: What advice would you give someone starting to code with AI today?
Doc Midnight: You know, it's funny. I had a young, quote-unquote, apprentice, somebody I was mentoring, and he asked me, "Is it even worthwhile learning to code at all, or can I just get away with what I'm doing?"
When I learnt to program, it wasn't so much about learning how to handle integers, floats, and things like that. It was about learning the stress tolerance required to fix something that's broken. It was always about learning how to learn, and learning how to find answers.
So what I told him was: you should still learn to code, and use AI while you do it. You'll get through it at eight times the pace, but you'll have the solid software engineering background you need to optimise it.
A lot of my security engineer friends were never engineers in their own right. They were very focused on vibe coding their way through apps, and that's caused all sorts of dramas for them, yet they've still got production systems running on that code. Meanwhile, a lot of my friends who came from an engineering background, where you learn good principles and how to write good tests, have gone very far, and most of them code with AI as well. So I think the proof's in the pudding there.
Tom: Just wrapping up: what are you excited about that's coming up, whether it's your own repos or things you're seeing out there?
Doc Midnight: There are a few things. Orro is the product we'll be looking at launching over the next three to six months. We'll be looking for security practitioners who'll actually use it. But because of the nature of the tool and the models under it, it can only go to people who've been vetted. It's just a responsibility thing. I don't think it's ethical to open-source that specific harness because of what I've built it to do. So I'm looking forward to launching it through a trusted pipeline shortly.
As for what's coming out: there are a lot of advancements in running AI models on local hardware and edge devices. That's become really critical for a few of my use cases, and for stopping the reliance on cloud, which I'm very big on. There are a lot of research papers coming out on how NAF‘s is working to do that, so I'm very interested in that. I'm also keen to use NVIDIA Pair across my stack. I've been wanting to play with it for a while but haven't had time since it came out. I think it's only two weeks old, but it feels like everyone's talking about it.
On repos: I've closed a lot of mine for the moment because work requires it, but I'll likely be opening them back up because of recent policy shifts that mean I can host them again. Once they're open, I've got a hell of a lot of different ones: around security, around development on the Microsoft suite like Power Platform, and automation in Azure from all my sysadmin days. I really enjoy Terraform for resource provisioning, so I've got a lot of stuff around that I use in my home stack.
I'll also be open-sourcing parts of Orro, specifically the continuous improvement harness, because I don't think it's right to keep that to myself. It can't be damaging unless you're already highly capable at something damaging. So yeah, I'll be open-sourcing parts of Orro and parts of the other things I've built over the years.
Tom: This was a great chat. Thanks, Doc.
You made to the end, sign up for the next AI pilled interview!
Once a month you'll join us and get deep into the weeds of the latest AI engineering with real practitioners.
Get the next interview
Comments ()