• dghlsakjg an hour ago |
    For those looking for the answer that don’t want to click through the entire repo: no. Kimi cannot do your American taxes, but origin doesn’t seem to matter.

    It appears that no LLM can get more than about a third of returns correct. Kimi K3 is on par with Sonnet 5. They are both very bad at doing your taxes.

    • persedes 20 minutes ago |
      Glad to see a benchmark outside of programming tasks. Presumably this one was not part of the training and shows the models performance (or lack thereof) on knowledge tasks.
  • yieldcrv 7 minutes ago |
    I like these tests

    In 6 months an LLM will get it to 80%

    people will exclaim negatively “it was trained on the eval [to do this incredibly useful thing]”

    meanwhile a problem is solved