Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Happy to answer questions about the eval methodology, the backend findings, or anything in the repo. I'll be around.


Thanks to everyone for the great discussion! v0.7.0 is out now. It was in flight when this landed - changes tool error channel based on dogfooding observations with some larger models, eval re-run (numbers shift but within CI), and most importantly docs updated! I hope they're clearer now.


v0.7.1 is now out - Forge can now sit behind Claude Code! Proxy mode can talk to supported backends and handles format translation. Anthropic > forge > OpenAI; or Anthropic > forge > Anthropic.

Will get this ported over into vLLM work and try to get that released soon.

Thanks to some kind folks who contributed Docker, token counting, and a handful of PRs I haven't gotten to yet.


been testing forge with ternary bonsai 8b mlx 2bit, pretty sweet even if the model is limited - real potential with this project, good luck!!

  - Broad slice:
      - Full Forge: 48/72 accurate, 72/72 complete, score 66.7%
      - Bare: 18/72 accurate, 24/72 complete, score 25.0%
      - Lift: +30 correct runs, no paired regressions
      - Bare had 42 ToolCallErrors and 6 ToolExecutionErrors; full Forge had none.
  - Advanced reasoning:
      - Full Forge: 3/24 accurate, 24/24 complete, score 12.5%
      - Bare: 3/24 accurate, 9/24 complete, score 12.5%
      - Lift: completion improved, but accuracy did not.


super interesting work. It will take me a few days to dig in and really understand it. But I'm looking forward to it.

I run small models at home, so I'm very curious.


That's awesome! Let me know if quick start is causing issues or anything else you'd like to dig into.

Out of curiosity, what models are you running?


dashboard link is dead



yes, that link works for me.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: