What I Wish I’d Known About Scaling Before Building with AI

BlackArbs Admin

Slow backtests, blown memory, and the coding habits I wish I’d caught sooner.

Man, so many lessons. I’ll get right into it. You start with an idea, a hypothesis, an anomaly or whatever and you want to test it. You want to see if there is anything valuable there. In the past you had a few paths you could take. You could just grab a jupyter notebook load up some free Yahoo data and do some EDA as a starting point. If it was still good then you would move on to a backtest.

From there you had to decide how serious you wanted your backtest to be b/c at that time it wasn’t a lot of trustworthy backtesters available. Most people would take the short cut and run a vectorized backtest in pandas or something similar. It was fast and cheap and allowed for quick iteration and optimization.

The problem was you exchanged speed for realism. So you could start incrementally adding more realistic components until you realized that you needed an event based backtester. Then came quantopian, quantconnect and others. They setup the platform and gave access to high quality data and a high quality platform to build and experiment with.

Now the friction there was you had to learn their platforms, and limitations, and moreover they were a business and they wanted to get paid. Serving tens of thousands of users high quality data quickly and robustly is a major cost so you had to pay.

Fast forward now we have LLMs. If you have the data, money for a subscription and time you can roll your own from scratch. I started doing exactly this and experimenting with them in fall 2024 with Claude sonnet 3.5. They were way buggier back then than now (and less restricted). Now they have improved massively and still they cannot be relied upon to build accurate trustworthy systems autonomously, yet. Anyway, that's a tangent I’ll dive into at a different time. However as a result of working with these tools I learned a lot and here are some things that I would tell myself if I had to start again.

1. Think about scale from day 1.

Old rules about premature optimization throw that shit in the garbage. Now that you can build quickly and “effortlessly” you will naturally want to solve more problems, run more data, more analysis, more backtests more everything. One of my biggest frustrations was starting small but correct, then wanting to scale and realizing I would have to essentially refactor and reengineer the entire project from the ground up due to the default antipatterns LLMs use when coding.

2. Include profiling and monitoring in your systems from the beginning.

When you want to scale and your previously sub couple minute “correct” “reduced real” backtest is now taking fucking hours to run, your LLM won’t save you. If you ask it without data it will guess at likely bottlenecks and be correct about obvious ones. But when you refactor it and the time savings are negligible you will have to add profiling anyways. Bonus kick in the balls is that if your system is moderately complex at all, this side quest becomes its own iterative project where you are trying to account for where all the fucking time is spent in your system.

3. Don’t just profile time, profile memory and consider your hardware resources early.

When I was a python jockey memory was hidden from me. Systems were much simpler, scale wasn’t a big deal, and a 16GB RAM laptop was a lot. Now it’s fucking scraps. LLMs by default will happily brick your system with OOM or CPU thrashing in an attempt to meet your request to scale. If it’s a personal computer it will lock up. If it’s a cloud computer expect huge fucking bills for no reason. You will think you need more capacity meanwhile the LLM is doing stupid shit like materializing and rescanning huge fucking chunks of data in unnecessary nested loops.

4. Learn the basic antipatterns that most LLMs will put in your code

LLMs will inject garbage into your system because they are trained on a sea of dogshit code they stole from stackoverflow and github. Here are a couple to get you warmed up:

  • Repeated full data scans, like recalculating a metric on every event using full history or data
  • I/O inside hot loops, fetching or saving data repeatedly in the most performance sensitive part of your code
  • Duplicated calculations/work etc. inside hot loops or in general. You need the feature dataframe to be passed to multiple consumers, instead of computing it once, let’s do it 2x or 3x times!
  • Unbounded materialization, see above. An example I can think of is loading huge chunks of data just to access a small chunk of data, repeatedly, without releasing or deleting the object. OOM hell lives here.
  • Repeated data/type object conversion. Example is going from a dict like object to json to a model back to a dict like object. Most of the time the conversions are totally unnecessary.
  • Bonus one, LLMs love json. Json is ok for small use cases but they have no problem using it like a mini unstructured database making everything that has to interact with that object slow and heavy.

That’s all that I could think of at the moment. I know there’s more but I’ll save that for a future update. What about you, have you encountered any of these issues in your own work?

Enjoyed this post?

Subscribe for more research and trading insights.

By clicking "Subscribe," you agree to our Terms of Use and acknowledge our Privacy Policy. You can unsubscribe at any time.

No spam. Unsubscribe anytime.