Am I wrong in thinking that scrapers are the real secret sauce tech here? I mean the state of the art transformer LLM tech is all based on the same public research. But creating a robust corpus of data...are scrapers totally mundane solved commodity tech in 2023?
I'd rather say the corpus itself. Of course scrapers might play a role in creating that corpus. But it’s just one part: maybe also OCR and OCR correction, format conversion (think of double column PDFs).