Why DuckDB 2.0 is faster

(motherduck.com)

77 points | by tosh 3 hours ago

4 comments

  • scythmic_waves 1 hour ago
    I also love the visualizations but I'm getting heavy LLM vibes from the prose:

    > One setting drives this,...

    > The cost is now about the rows you actually touch, not rounds times table size.

    Etc.

    I get the brain scramblies [1] from trying to parse this writing style at work so I hate to see it elsewhere. Apologies if I'm wrong. But if I'm not then OP don't use an LLM to write for you. It's hazardous to your reader's health [2].

    [1]: https://www.youtube.com/watch?v=ipUJq-odt5Q

    [2]: https://discourse.haskell.org/t/how-to-keep-enjoying-program...

    • kristianp 26 minutes ago
      I noticed some AI tells, but found overall the article not too bad. It did seem to waffle at times though.

      > Storing a lake as thousands of 1 MB Parquet files is a bad practice anyway, and 2.0 does not rescue it.

      The "does not rescue it". No human would write like that.

      > I'll explain what that means on a table you already know.

      No I don't already know that table.

      Also

      > and claims 40x on graph reachability

      Is really hard to parse.

      The section on recursive CTEs wasn't well written and didn't explain how the optimisation was done. This article explains how the recursive CTEs were improved https://duckdb.org/2026/08/25/how-duckdb-runs-recursive-ctes...

    • majormajor 17 minutes ago
      I've lost count recently of how many times Claude has given me something in this style that I can't understand, and then I ask it to rework parts of it, and then it tells me that the original things it claimed weren't actually quite right anyway.

      I'm starting to read it as a sign of low LLM effort not just low human effort. It seems most common when one few-sentence prompt leads it to generate 4+ paragraphs (and the longer the output, the worse the odds). Prompting to dig into each resulting paragraph one by one, to make them readable, makes it do higher-effort deep dives.

    • sagarm 54 minutes ago
      It really is unreadable. I guess you're supposed to skim an AI summary. Too bad you'd never see the visualizations that way.
    • hobofan 18 minutes ago
      > It's hazardous to your reader's health [2].

      Just because some guy on some forum said that doesn't make it true. That's not how you establish facts regarding health claims.

    • gbalduzzi 32 minutes ago
      I may be in the minority but I didn't get LLM vibes from this. Not enough to be bothered by it, at least
    • Eridrus 33 minutes ago
      I have tried out the alpha releases of 2.0 and I got slopcoded vibes from it.
    • j-pb 1 hour ago
      [flagged]
      • meerita 1 hour ago
        [flagged]
        • status_quo69 1 hour ago
          Where were the facts invalidated? Why can't a reader despise a writing style that LLMs converge towards?

          I think you're right on your last point, btw, nobody will care and we'll be the poorer for it. Just like how McDonald's and Walmart have driven out and undercut localized taste and culture in the US, so it shall be with writing of all kinds. It's happening right now and readers are getting used to reading the intellectual equivalent of a Big Mac.

          • gwerbin 1 hour ago
            It's not even a matter of despising or liking the style, it's a matter of actually being able to interpret it and makes sense of it in a useful way. Even Claude sometimes gets confused by the dense metaphor-heavy obtuse Claude style.
            • status_quo69 1 hour ago
              True, but I don't know. I'm not going to gatekeep the transmission of information (how could anyone?) but there's a certain level of ennui here for me and the easiest tell comes from the writing style. I still think there would be something lost in "claude write me a blog, don't use load-bearing".

              We invented algorithms and machines that ostensibly understand and interpret our words and can make things happen, which is a revolution of understanding. From a legibility perspective, these machines now can "understand" me just as well as if I were writing like Gene Wolfe, yet instead of embracing the idea of "wow I can talk to a computer" we have an overweight on "wow, I can have computer talk to everyone else on my behalf!".

              Even if the transmission of information here was perfect I'd still question the utility of outsourcing your brain on writing a very short article compiled from release notes.

          • filcuk 1 hour ago
            [flagged]
        • scythmic_waves 1 hour ago
          At what point did I "invalidate facts"?
          • arcanemachiner 1 hour ago
            Do not engage. This thread is loaded with bicker-bait.

            Apparently the readability of LLM prose seems to trigger stupid "tabs vs. spaces" style arguments. I encourage you to save your time and energy and do something productive with it instead.

            • scythmic_waves 48 minutes ago
              Great advice, will do. I don't engage with HN a ton and I didn't realize this was the state of discourse on LLM prose.
            • lopatin 56 minutes ago
              Agree, bike shedding about something like writing style, which unlike tabs-vs-spaces, doesn't even have an objectively correct answer, is not productive.
    • augment_me 1 hour ago
      My manager was previously whining about something similar when it comes to generated reports and I just pulled down all of his writing and made the LLM write in his style, he is really content now. Do the same using your own writing and you wont have to be offended
    • gchamonlive 47 minutes ago

        I get the brain scramblies [1] from trying to parse this writing style at work so I hate to see it elsewhere
      
      This sounds to me a lot like mass hysteria, people reading other people's behaviour online and reproducing it unconsciously.

      More and more really important and useful information will arrive like this for us to consume. There's no way around. So this is a disservice for newcomers that could come and go unscathed but instead is crippled by these kinds of comments that brings nothing of substance to the table and has the potential to make them hate something they otherwise wouldn't even notice.

      Also, from https://news.ycombinator.com/newsguidelines.html:

        Please don't post shallow dismissals, especially of other people's work. A good critical comment teaches us something. 
      
        Please don't complain about tangential annoyances—e.g. article or website formats, name collisions, or back-button breakage. They're too common to be interesting.
      • majormajor 20 minutes ago
        Look at the specific AI habits called out in this comment: https://news.ycombinator.com/item?id=50037220

        This sort of writing decreases readability. People say "just have your own LLM rewrite it" like you don't lose value when you go from prompt -> slop you didn't review enough to clean this crap up in -> someone else's prompt to change the style -> finally someone reads it.

        Do you know what happens a double-digit percentage of the time when I ask Claude to rewrite some shit that it gives me like this? It says things like: "I overstated this, I rechecked and actually..." or "this claim doesn't hold up, actually [this other thing is true]..."

        So it's a sign that the claims in the post likely weren't vetted very hard.

        So if you aren't proofreading I'm gonna be skeptical. And saying "deal with it" doesn't rescue it. Does it?

        And then there's the reflexive "you must just be an ideological hater." No. I'm someone who uses the tools in a domain where quality matters enough that I have to dig into the quality of the tool output and spot the tells for when it's low output, so that I can deliver shit that works reliably and consistently.

        • gchamonlive 15 minutes ago
          This is all absurd discussion. If you disagree with the article style, there is a flag button there.

            So it's a sign that the claims in the post likely weren't vetted very hard.
          
          Our time is better used submitting something useful instead of debating meaningless guesswork in the comments.
      • johnfn 38 minutes ago
        Curious how you see low-effort writing as a "shallow dismissal" or "tangential"
        • gchamonlive 30 minutes ago
          Everything is low-effort writing if your judgement is ideologically charged. Just read the damn post, it's very good, there's nothing low effort, but you need to read it as one single piece, not dissect it looking for hints of LLMisms, killing the article in the process.

          Also if you think LLMisms so bad it's spam, you also don't get a free pass. From the guidelines:

            If a story is spam or off-topic, flag it. Don't feed egregious comments by replying; flag them instead. If you flag, please don't also comment that you did.
          • johnfn 6 minutes ago
            Who said I was "ideologically charged"? I like AI and I use it every day. I simultaneously hate AI writing. It is impossible to parse. If you can read it fine, good on you, but you seem to fundamentally misunderstand that other people perceive the world differently than you do. No one is "dissecting" this article under a microscope, it's obvious.
          • qotgalaxy 8 minutes ago
            [dead]
  • stacktraceyo 2 hours ago
    Great visualization. Side note their new c++ extension api is also gonna be faster from the perspective of development / distribution of those extensions
  • jiggawatts 9 minutes ago
    I wish more database engines used a Task-based design like Umbra / CedarDB.

    Most of the DB engines out there still seem to use a "n-threads" style parallelism with exchange operations and poor async I/O management.

    DuckDB is improving on this front, but in some sense is catching up to R&D (and implementation!) that is now decades old.

    A "rhetorical challenge" I like to give software developers working on systems like this is the following: If I gave you a computer with 1,024 cores and matching network and storage bandwidth -- but with significant latency -- could you keep a system like this 100% utilised with one query?

    The answer for almost all software is "no".

    For example, SQL Server tops out at 64 hardware threads for any one query: https://learn.microsoft.com/en-us/sql/database-engine/config...

    GPU codes are starting to get there, but CPU codes are way behind on this frontier of computer science.

    It's not just databases! Can you (de)compress a file in parallel? Verify its hash in parallel? Upload/download from storage with CPU and I/O task parallelism? Can you overlap all of these operation so nothing is ever waiting on anything else it doesn't have to?

    This matters! I ran some tests with bioinformatics codes and found that they all got stuck in tar pits. None could scale to modern SSDs with millions of IOPS or modern networking with hundreds of gigabits of throughput, no matter how many CPU cores were thrown at them.

    PS: AMD's Zen 6 era EPYC 9006 processors will have 512 cores and 1,024 threads per two-socket system, so this is not hypothetical: https://www.amd.com/en/products/processors/server/epyc/9006-...

  • larodi 1 hour ago
    ...because Atlas and Fable took turns to go back and forth through it (source code and runtime) and track suspected bottlenecks.
    • gwerbin 1 hour ago
      Totally valid technique in my opinion. Even with last year's technology, LLMs were really good at synthesis of small details, which, combined with their encyclopedic knowledge of everything ever written down about computers and programming, and their infinite capacity to run adhoc experiments and build out test infrastructure, made them very good at debugging and performance optimization.