Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

No, it does not work like that, it actually can process the image itself there is not an intermediate image to text step


How does a Large Language Model process images then?


It can only deal in tokens, so you're essentially right that it creates a textual description before describing it back to you. This process is obviously incredibly lossy and details are easily missed




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: