> Spent over $400,000 in API priced tokens with GPT-5.6 Sol and GPT 6 Astra
Tangent here, but I think this bit is super interesting!
You could viably hire someone to do this work for that kind of money - I think the interesting thing is that substantially less interested/experimenting engineers would consider paying for a human to do this work, than would happily chuck a big amount of money into an LLM.
I don't have any suggestion about why that exists, but it's a strange and interesting contract.
I found the whole section super interesting... I'll copy it here:
I used a lot of OpenAI models to try and complete this port. In total I did over $400,000 in API priced tokens with GPT-5.6 Sol and GPT 6 Astra. They wrote over 1.3m lines of Rust over multiple months of /goal loops and never got past like 84% compat.
When I saw how little my Claude Code limits were burning, I figured it'd be fun to throw Opus 5.5 at this. It had a working v0 in 10 hours.
I assumed it kept using the code the Codex models wrote. I was wrong. Opus 5.5 started from scratch. It got further than Astra in 1/10th the time.
I let it keep going, and it definitely did. Total token spend was ~$24,047 of API spend over 2 weeks. I was using my Claude accounts, and it worked out to somewhere between 925% and 983% of my $200 plan weekly limits.
Expensive, for sure, but not that bad considering how much work has went into typescript-go.
> I assumed it kept using the code the Codex models wrote. I was wrong. Opus 5.5 started from scratch. It got further than Astra in 1/10th the time.
Ouch, that is brutal and honestly, quite embarrassing but confirms what I have been seeing for a while. Personally, I find output from current OpenAI models still very hard to parse (though it has gotten better vs the pre-trains from both labs in mid/late 2025), thus hard to truly understand, verify and get comfortable maintaining vs current Anthropic models. I do occasionally see a higher ceiling in well scoped tasks with OpenAI models at the cost of (frequently) deviating from the original prompt in (sometimes) very destructive ways.
Could be that this hard-to-read output doesn't just go over my limited capacity/skills but with current models can become simply impossible to untangle beyond a certain size even when one has (essentially) infinite resources via multiple subs and different models.
Would also work with my suspicions for why OpenClaw (mainly build with Opus 4.5 and its post-trains) has been this hard to truly "fix", requiring highly paid Nvidia engineers, multiple months, (literally) infinite resources from OpenAI including access to internal models and yet still holds records for CVEs. Heck, another one was found just 7 days ago after what I'd argue was one of the most extensive hardening sessions any piece of software has ever undergone.
Makes my (multiple) decisions to start from scratch more than once on a major reworking of the existing tabbing interface in Firefox a bit less painful. Learned with each, found gaps in my knowledge, thanked the amazing docs the Firefox devs have been maintaining for decades and while starting from 0 was painful, getting back to MVP is easier than ever. When I hit a point were I was starting to struggle to truly parse additions a model was making to the patches applied to Firefox source code (even if they worked), I always found that pushing even slightly beyond that would incur painful, but hard to notice regressions, introduce major DB maintenance burdens as some models struggle to understand that in development regressions and incompatibility are acceptable and schema transitions aren't needed pre-release (still a case with GPT-6.1 Sol, less so post Fable for Anthropic), make me uncomfortable concerning privacy/security/data loss prevention and simply take away my control about the implementation. Wouldn't feel right to release something in that state, what I have now is fully understandable and thus could be maintained even without models.
Still expecting bugs of course, massively dreading security findings or even worse, possible data loss given browsers handle some of our most important personal+professional data and will surely have taken some embarrassing approaches that might have a much more performant solutions when implementing an infinite canvas of webpages, but still, rather that then also knowing I wouldn't even know where to start understanding a feature.
More so if, like with ts-rust, even (nearly) infinite tokens couldn't get me unstuck.
This is Theo's project, and he has been very transparent about how he is only doing these things because api prices are heavily subsidized by subscription pricing. The reason people aren't willing to pay 400k to an engineer to do this work is because they aren't paying the AI this much.
I don't think I've seen an enterprise paying API prices doing these sorts of rewrites, and honestly, I think until that happens I will remain skeptical about LLM AI's viability as a profitable business.
>Spent over $400,000 in API priced tokens with GPT-5.6 Sol and GPT 6 Astra. They wrote over 1.3m lines of Rust over multiple months of /goal loops and never got past like 84% compat.
Then they let Claude code lose on the problem:
> I figured it'd be fun to throw Opus 5.5 at this. It had a working v0 in 10 hours. I assumed it kept using the code the Codex models wrote. I was wrong. Opus 5.5 started from scratch. It got further than Astra in 1/10th the time.
In the warnings from README.md
> Also worth mentioning: I've never read a line of this code.
TypeScript team member here - this is impressive work, and it's honestly crazy that 3 of these ports have popped up in the last week! I'm currently AFK but we're hoping to learn more and we'll have more to say on this soon.
I suppose my comment sort of implies this. I think the naive knee jerk reaction is that if you can spit out perfect machine level code against a specification, then logically being closer to bare metal would help with performance. However, without knowing the details of how things worked it is tough to say whether it is truly "reasonable" or not.
But more importantly, the Typescript team themselves [1] picked Golang mostly out of ergonomics and ease of porting at the time. It is becoming more increasingly more to "taste" what ergonomics means. Notably, they did not pick it because they believed it to be the fastest option. But speed was on the mind, giving way to ergonomics and ease of porting. At least, this is my read.
The notable thing that these LLM-generated ports aren't doing is _rewriting in Go_. TypeScript 7 is a lot faster than the JS/TS-based version 6, but memory use is a major bottleneck.
However, it mostly represents a mechanical port of the old codebase to Go. What I haven't seen anyone do is try a true rewrite in Go optimized for performance (whilst keeping a strong emphasis on readability).
Very interesting. With rust's ocaml heritage, I'd imagine a straight port from typescript would be simpler and more idiomatic than a port to go. I guess not.
Language and compiler design/ implementation is very much a human task. There's lots of nuances to consider. I doubt they want agent spaghetti code in their codebase either.
GC doesn’t affect performance (at least visibly), memory allocations on critical paths do. Moreover, Rust ARC may cause memory fragmentation and perform worse than a GC.
Writing Rust, Zig, or Go, you would still control memory manually, where it matters.
GC absolutely has performance costs, even with minimal allocations, because tracing collectors must scan live objects. This is exactly what this post says. Don't know if its really improved over time in real world cases.
It does have an impact. Discord blogged about unavoidable GC delays in go that occurred because go needed to traverse live memory, their only option (in go) for reducing that pause was to reduce the size of their data set.
It depends on the workload. ARC can be more expensive than tracing GC, especially with heavily shared objects.
But my main point was that the presence of a GC has nothing to do with how close a language is to the metal. Memory management strategy and low-level capabilities are two separate things.
I’m not saying it doesn't work. I’m saying it’s a waste of resources and won’t gain widespread popularity. The current trend is porting source code to languages that generate native code.
I dislike the WebAssembly advocacy of some folks, where they act as if there weren't similar attempts all the way back to UNCOL in 1958, and even in recent history, seem to only focus on JVM, while ignoring the CLR and BEAM.
However, there are indeed some cases where it does make sense since we have it around anyway, and one of them is running the Typescript compiler in the browser, now that it is written in a compiled language.
It's extremely useful if your distribution platform is the web. However, I find it much more valuable as a type checker than a compiler. I use TypeScript extensively (in several projects) for defining schemas types.
For example, in the Breaka Club (https://breaka.club/) editor, which is not presently exposed to kids yet, we define behaviors for characters, props, mosaics (terrain tiles) and items in JSON. But the JSON has type checking and real-time completion as you type. Not naive property auto-suggestion, we have effect types and they're context aware in that you can refer to targets introduced higher in parent constructs of the JSON — and they're fully type checked. If you refer to a target that does not exist in a context, it won't validate and you cannot save the schema.
We use tsgo for development, but the editor (runtime schema validation) is presently stuck on TypeScript 6 because there's no WASM support.
Now, is it nutty that we're using TypeScript for JSON validation? A little. But it's extremely powerful. We're going far beyond what's capable with Zod or ArkType.
Theo is making a new twist on the idea of IDE which hosts the application itself. I'm sure there's going to be a video about it soon. But the idea is to embed the whole compilation and type-check in wasm in the browser which then runs the app immediately, as the agents edit the code.
Wasm is a first-class target for Rust, which produces faster and smaller Wasm binaries. My understanding is that you need to use TinyGo to reasonably target Wasm with Go.
Go having a garbage collector isn't necessarily bad, since Wasm has GC, but I think Go's GC can't easily use Wasm GC because Go has interior pointers.
I'm assuming they mean because it doesn't need a garbage collector or much of a runtime at all. You can target direct to WASM and use I guess WASI, while Go brings with it a garbage collector and so on.
Was it ever a reasonable choice? At the time the choice was made the tooling ecosystem was already split between Rust and JS, and by putting TS in Go they chose to split it three ways. Why not split it 4 or 5 ways then? The pressure will always be towards less duplicated work.
The only real choice of language to build the next generation of JS tools in is JS. Anything else is a vote of no confidence in ourselves.
It's ok for a language (JS in this instance) to be good in a domain but not good in all domains. Pulling a language in every direction forces it to make compromises that hurt its applicability in specific areas.
I think the future of this is -- get your AI agent to investigate this port for anything actually worthwhile, comparing it to the go version, and then take those good ideas and actually maybe bring them into the real compiler.
What a waste of power... $424.000 of API priced tokens amounts to how many kWh burned? And how would you trust this to build your code?
Motivations says it all:
Motivations
Test model capabilities
Make a fast TypeScript type checker
Make a ts checker that can work in WASM with high performance
Memes
Maybe there are just so many of these but at this point for vibe coded stuff like this I’m just like “who cares.” LLMs can make stuff like this now. Unless this reaches some critical mass of usage among developers why would I use it? It’s just a less maintained, less tested implementation. The more important question is “what does this teach us?” And idk the answer to that one.
Yes. This has nothing to do with people trying to rewrite things in Rust for fun. It's a part of his development project that I'm sure will get a video in the future.
I think the difference here is that the original motivation for choosing Go over Rust was that Go was easier to port TypeScript to.
LLMs change the calculus a lot, the result seems faster, and it's possible that the TypeScript team might want to change directions. It'd be a big deal, but I bet they'll at least discuss it.
Generally building something is not about "What does this teach us?" but rather "What can this be used for?". You don't build a house and as "What does this teach us?". It's honestly almost a perverse question to primarily ask for something that aims to have practical value.
The clear reason this has worth is that it's much faster than TSC.
HN is going to be a weird place to be between people loving AI and the loss of loving tech. I don’t remember last time I watched a good tech dev video. At least the video game industry is a little bit shielded from this but soon it will be the same when we see the speed AI is advancing.
I don't see at all that this is stifling love of tech. Sure if you liked to engage in language wars or vim vs emacs or tabs vs spaces then yes that's over. But tech as we know it is currently exploding, a ton of stuff that simply was not practically possible before is now possible.
The point is that it's clearly not just a less maintained, less tested implementation. It's 1.61x faster for one in real world apps, in some case even almost 4x faster.
I am curious why it's faster. Golang isn't that much slower than Rust, aside from the garbage collector. 4x performance improvements is much bigger than what I expected.
This was a lot of tokens to spend on a port. I wonder if all of the porting activity out there would be better supported by investing in really good deterministic automatic translators that do 90% of the job cheaply and delegate any decisions that have to be made back out to the AI agent?
I have to wonder if LLM reimplementations verified against years of human tests will have holes where original authors thought that tests are not necessary and common sense is enough.
I have to wonder this to preserve my ego as a human. I wonder have to this because I sure as hell am never going to go through all that code to check.
Of course they will. For about 3 days, and then it will get fixed.
At this point, for software which has an extensive test suite (or a large body of interoperable software), humans are more likely to create these kinds of gaps than the LLMs are.
LLMs have been trained to fix compile errors in a loop. Give them years of human labor worth of tests and they will fix compile errors until it works (as well as the tests can ensure). Now you have a codebase no one has read. Good luck adding new code to it.
Let's also not forget that Bun was claiming a rewrite in 11 days or whatever it was, but actually spent three months of human labor fixing hundreds of issues their rewrite introduced before shipping it as a release.
Probably zero, or even a profit. The idea that these companies are making a loss on subscriptions is an unsubstantiated HN fallacy.
On the contrary, they're probably making a small profit on subscriptions, or at least break even, and absolute bank on API pricing. Anthropic recently reported an 80% gross profit margin, for example.
It might, but also might not. I would imagine the economics are genuinely very tangled up even before you start obfuscating it - it’s very hard to fairly allocate fixed costs between training and inference.
I think a clearer picture would be from inference-only orgs.
Sibling dead comment said it already but that's likely an estimate of what it would cost via API and in practice it was a few hundred or thousand dollars of subscription fees.
I'm curious how much unsafe this uses? Are these ports making idiomatic code or just using unsafe all over. There doesn't seem to be anything in the readme.
As much as possible mechanical (non-LLM) verification of equivalence that generates good diagnostic messages, mechanical translation steps that generate good diagnostic messages, a good memory system to reduce rework... (edit: and mechanical coverage measurement)
It is beginning to look like the future is treating Rust as an optimization pass. Write your program in a higher-level language like Typescript, Go, Python, etc. and then have LLMs compile that to Rust.
Not really, the future are natural languages, with some markdown files for formal specification, yaml or whatever, and then Assembly gets generated directly.
3 GLs are going to slowly fade out as LLM tooling improves, and coding looks increasingly as envisioned by 4 GLs advocates.
Sounds awful, natural language sucks for describing complex problems and relationships. That's why people have historically reached for diagrams and pseudo code.
Tangent here, but I think this bit is super interesting!
You could viably hire someone to do this work for that kind of money - I think the interesting thing is that substantially less interested/experimenting engineers would consider paying for a human to do this work, than would happily chuck a big amount of money into an LLM.
I don't have any suggestion about why that exists, but it's a strange and interesting contract.
I used a lot of OpenAI models to try and complete this port. In total I did over $400,000 in API priced tokens with GPT-5.6 Sol and GPT 6 Astra. They wrote over 1.3m lines of Rust over multiple months of /goal loops and never got past like 84% compat.
When I saw how little my Claude Code limits were burning, I figured it'd be fun to throw Opus 5.5 at this. It had a working v0 in 10 hours.
I assumed it kept using the code the Codex models wrote. I was wrong. Opus 5.5 started from scratch. It got further than Astra in 1/10th the time.
I let it keep going, and it definitely did. Total token spend was ~$24,047 of API spend over 2 weeks. I was using my Claude accounts, and it worked out to somewhere between 925% and 983% of my $200 plan weekly limits.
Expensive, for sure, but not that bad considering how much work has went into typescript-go.
Ouch, that is brutal and honestly, quite embarrassing but confirms what I have been seeing for a while. Personally, I find output from current OpenAI models still very hard to parse (though it has gotten better vs the pre-trains from both labs in mid/late 2025), thus hard to truly understand, verify and get comfortable maintaining vs current Anthropic models. I do occasionally see a higher ceiling in well scoped tasks with OpenAI models at the cost of (frequently) deviating from the original prompt in (sometimes) very destructive ways.
Could be that this hard-to-read output doesn't just go over my limited capacity/skills but with current models can become simply impossible to untangle beyond a certain size even when one has (essentially) infinite resources via multiple subs and different models.
Would also work with my suspicions for why OpenClaw (mainly build with Opus 4.5 and its post-trains) has been this hard to truly "fix", requiring highly paid Nvidia engineers, multiple months, (literally) infinite resources from OpenAI including access to internal models and yet still holds records for CVEs. Heck, another one was found just 7 days ago after what I'd argue was one of the most extensive hardening sessions any piece of software has ever undergone.
Makes my (multiple) decisions to start from scratch more than once on a major reworking of the existing tabbing interface in Firefox a bit less painful. Learned with each, found gaps in my knowledge, thanked the amazing docs the Firefox devs have been maintaining for decades and while starting from 0 was painful, getting back to MVP is easier than ever. When I hit a point were I was starting to struggle to truly parse additions a model was making to the patches applied to Firefox source code (even if they worked), I always found that pushing even slightly beyond that would incur painful, but hard to notice regressions, introduce major DB maintenance burdens as some models struggle to understand that in development regressions and incompatibility are acceptable and schema transitions aren't needed pre-release (still a case with GPT-6.1 Sol, less so post Fable for Anthropic), make me uncomfortable concerning privacy/security/data loss prevention and simply take away my control about the implementation. Wouldn't feel right to release something in that state, what I have now is fully understandable and thus could be maintained even without models.
Still expecting bugs of course, massively dreading security findings or even worse, possible data loss given browsers handle some of our most important personal+professional data and will surely have taken some embarrassing approaches that might have a much more performant solutions when implementing an infinite canvas of webpages, but still, rather that then also knowing I wouldn't even know where to start understanding a feature.
More so if, like with ts-rust, even (nearly) infinite tokens couldn't get me unstuck.
I don't think I've seen an enterprise paying API prices doing these sorts of rewrites, and honestly, I think until that happens I will remain skeptical about LLM AI's viability as a profitable business.
>Spent over $400,000 in API priced tokens with GPT-5.6 Sol and GPT 6 Astra. They wrote over 1.3m lines of Rust over multiple months of /goal loops and never got past like 84% compat.
Then they let Claude code lose on the problem:
> I figured it'd be fun to throw Opus 5.5 at this. It had a working v0 in 10 hours. I assumed it kept using the code the Codex models wrote. I was wrong. Opus 5.5 started from scratch. It got further than Astra in 1/10th the time.
In the warnings from README.md
> Also worth mentioning: I've never read a line of this code.
I am so confused about the motivation behind this.
Seems obvious, experimenting with what it's possible with LLMs.
But more importantly, the Typescript team themselves [1] picked Golang mostly out of ergonomics and ease of porting at the time. It is becoming more increasingly more to "taste" what ergonomics means. Notably, they did not pick it because they believed it to be the fastest option. But speed was on the mind, giving way to ergonomics and ease of porting. At least, this is my read.
[1] https://github.com/microsoft/typescript-go/discussions/411
However, it mostly represents a mechanical port of the old codebase to Go. What I haven't seen anyone do is try a true rewrite in Go optimized for performance (whilst keeping a strong emphasis on readability).
https://news.ycombinator.com/item?id=22336284
It is one person and one project, but I found it interesting to read about his experience.
I like Go, but I probably would have tried Rust first because I want pattern matching when implementing languages.
Writing Rust, Zig, or Go, you would still control memory manually, where it matters.
GC absolutely has performance costs, even with minimal allocations, because tracing collectors must scan live objects. This is exactly what this post says. Don't know if its really improved over time in real world cases.
But my main point was that the presence of a GC has nothing to do with how close a language is to the metal. Memory management strategy and low-level capabilities are two separate things.
However, there are indeed some cases where it does make sense since we have it around anyway, and one of them is running the Typescript compiler in the browser, now that it is written in a compiled language.
For example, in the Breaka Club (https://breaka.club/) editor, which is not presently exposed to kids yet, we define behaviors for characters, props, mosaics (terrain tiles) and items in JSON. But the JSON has type checking and real-time completion as you type. Not naive property auto-suggestion, we have effect types and they're context aware in that you can refer to targets introduced higher in parent constructs of the JSON — and they're fully type checked. If you refer to a target that does not exist in a context, it won't validate and you cannot save the schema.
We use tsgo for development, but the editor (runtime schema validation) is presently stuck on TypeScript 6 because there's no WASM support.
Now, is it nutty that we're using TypeScript for JSON validation? A little. But it's extremely powerful. We're going far beyond what's capable with Zod or ArkType.
- https://www.typescriptlang.org
- https://code.visualstudio.com/docs/remote/vscode-web
Go having a garbage collector isn't necessarily bad, since Wasm has GC, but I think Go's GC can't easily use Wasm GC because Go has interior pointers.
However, it has a relatively large runtime, so TinyGo is often preferred when bundle size needs to be small.
The only real choice of language to build the next generation of JS tools in is JS. Anything else is a vote of no confidence in ourselves.
Maybe. Who knows, it's all changing so fast!
Motivations
LLMs change the calculus a lot, the result seems faster, and it's possible that the TypeScript team might want to change directions. It'd be a big deal, but I bet they'll at least discuss it.
The clear reason this has worth is that it's much faster than TSC.
No more joy crafting things as a dev…!
> It’s just a less maintained, less tested implementation
Then they asked if it teaches something because it is a risky utility to them. You're trying to make it sound a lot weirder.
I already pay for subs on both but the tokens are already spoken for with other projects
I have to wonder this to preserve my ego as a human. I wonder have to this because I sure as hell am never going to go through all that code to check.
At this point, for software which has an extensive test suite (or a large body of interoperable software), humans are more likely to create these kinds of gaps than the LLMs are.
Let's also not forget that Bun was claiming a rewrite in 11 days or whatever it was, but actually spent three months of human labor fixing hundreds of issues their rewrite introduced before shipping it as a release.
wut?
I personally find it ridiculous. The cost is what you paid, not what someone else could have paid. Markets and all that.
It would be like using spot instances on AWS but flexing by boasting about how much on-demand $$ compute you used.
It's also TheoGG sooooo.. Those influencer instincts tho.
I wonder how much AI companies lost on this?
Edit: not that I’m worried about them, I’m just curious
Probably zero, or even a profit. The idea that these companies are making a loss on subscriptions is an unsubstantiated HN fallacy.
On the contrary, they're probably making a small profit on subscriptions, or at least break even, and absolute bank on API pricing. Anthropic recently reported an 80% gross profit margin, for example.
https://www.reuters.com/business/retail-consumer/anthropic-t...
I think a clearer picture would be from inference-only orgs.
i'd bet this is a few subscriptions over a few months rather than paying directly for tokens.
Doesn't look like any on skim, also looks like an explicit goal was to avoid unsafe code: https://github.com/pingdotgg/ts-rust/blob/ad8f2746ea85a6354b...
3 GLs are going to slowly fade out as LLM tooling improves, and coding looks increasingly as envisioned by 4 GLs advocates.
Is this going to be maintained? Nope
Nevertheless this and many others are showing what is possible. Much akin to how someone ports doom onto a microwave.
We'll look back at this time with great fondness when we collectively discovered a whole new gear.
More unserious and unmaintained projects at 400k a pop. Sick.