I have multiple LLM subscriptions at any given time, plus an array of local models.
When I ask a question outside of my domain of expertise I like to ask all of the LLMs I have access to. I also create separate sessions and ask the same question multiple ways.
It’s revealing to see how many different and contradictory answers I get, most of which are presented confidently.
The last time I ran a medical question through Claude I couldn’t even get consistent answers between sessions.
It’s also scary how easily you can lead each LLM to the answer you have in mind. When I would start asking questions about different options that other LLMs had presented, each session would drift toward that explanation.
In my day job we tried creating a credit assessor tool using LLM as the credit assessor.
It did great, generated a report on the assessed business that was incredibly detailed and plausible.
Then I started running tests and getting into the details, and found that if you ran the same report on the same data, it generated completely different, still very plausible, results. I could run the same source data through the assessment process 10 times and get 10 very different results. We had to can the project and go a different route.
LLMs are designed to produce plausible results, not factual results. We can fix this when using them for software dev by using linters and tests (though we've all had the experience where the LLM invents an API endpoint). I would not trust raw LLM output in any situation where that kind of testing and verification capability isn't present.
What's crazy is that there are ton of businesses building processes around LLMs that haven't done this exercise and fully believe the LLM is giving them accurate data.
> LLMs are designed to produce plausible results, not factual results.
They are true to their name: Language models. It is precisely the same problem in a language: a grammatically correct sentence is not necessarily true.
Yup I use llm to write scripts for me to process data I don't ask the llm to process the data themselves. Even when I wrote something for my day trading I used llm write scripts that do all the processing and predict price movement from that the more data is pre processed the more all the llm come up with similar trades.
It's funny that if the LLMs had all given the same result each time (it sounds like) you would have considered it more valid, even though it might just be giving a single wrong answer more consistently.
if two llms independently cite the same law given a certain factual legal question then it generally makes it more likely that the law cited does, in fact, apply to the fact pattern in question ; i think you’re using “right” in a moral sense irrelevant to the comment i was responding to.
your follow-up makes no difference to the main point under discussion : if you ask two different llms about a legal fact pattern and they independently reference the same law, then it is, in fact, more likely that the law is more relevant to the fact pattern than if only one llm mentioned it.
My point was that a human can say "sure laws say this but how about it's a screwed up thing and we change it for the better" and LLM will just do whatever is the highest in probability space
yeah, it would have been harder to see that the LLM was making shit up if it had consistently made the same shit up. But this is also true of humans, so...
you can set the "temperature" which is a lever on how stochastic the prediction is. If you are doing your own inference this is clear and easy. If you are consuming tokens this is outsourced.
What happened to VERIFYING an answer? Does nobody do that anymore?
When I ask an LLM, I trace the sources, and see if they make sense.
More often than not the sources don't actually say anything about the topic in particular...
> It’s also scary how easily you can lead each LLM to the answer you have in mind.
Exactly. Which is why "treat an LLM like a human expert who can answer your question" doesn't work. It's more like a human bullshitter who makes up convincing looking answers, and tries to please you. If the answers have actually some grounding in the training material, that's useful as some kind of holistic google, but often it's not.
> What happened to VERIFYING an answer? Does nobody do that anymore?
The problem with medical advice is that you may not be competent to verify the answer, right?
I agree that asking 5 LLMs to vote and trusting the answer is totally the wrong approach, of course. But LLMs (and traditional material) can help getting more informed. For instance, instead of going to your doctor with the LLM diagnosis and trying to convince the doctor that the LLM is right, you can try to build your own understanding of the problem and go ask the doctor to explain to you what you understood correctly and what you misunderstood.
If you have some understanding, it's harder for a specialist to bullshit you. But you need your own critical thinking and you need to put effort into actually learning something, blindly trusting and repeating what LLMs say doesn't help.
Or more specifically in this case: the patient was obviously insisting on a diagnosis and treatment based on ... a slightly hurting shoulder, with zero visible or detectable phenomena.
So the doctors gave him what he wanted: a treatment ... and Claude told him the treatment was a placebo. Correctly, I might add.
Yeah, it is absolutely not what the patient wanted to hear. BAD doctors! Except ... no, not really.
Hmm that's not what I read from the article. The author says the opposite, actually:
> [the orthopedist] suggested I get an MRI, which the clinic conveniently had available. [...] This, of course, means little to me, but their suggested course of treatment was extensive; [...] Coming out of the clinic, I had the feeling they had jumped the gun.
The author also says:
> They injected me with Traumeel, which is registered in Germany as a homeopathic medicine "without a therapeutic indication".
I personally wouldn't want to be injected homeopathic medicine "without a therapeutic indication" without even knowing it is homeopathy.
And my recent experience with multiple doctors at multiple hospitals is that just like LLMs, you shouldn't blindly trust them. Sometimes they make mistakes, and in my experience they never, ever admit it (maybe even to themselves).
So trying to get informed "on the internet" (including with LLMs) feels sane to me, but that's worth what it is worth.
What concerns me in the article is that they have GPT make a diagnosis, Claude review it, and somehow seem to assume that if both LLMs agree, they must not be completely wrong. Just like for code, it takes an expert to leverage an LLM to take an expert decision. A beginner can leverage the LLM to understand the problem better, but they never reach the level of expert just from that.
Unfortunately it's part of lowering the confidence towards doctors, and the solution to that is to try and get informed and ask another doctor. And of course they don't like it if you say "I already asked someone else but I don't trust doctors in general, so I am now asking you to test the both of you".
Quite the opposite. The doctors are disagreeing with the LLM, and when the patient doesn't accept, they're putting in a minimal lie in the deontologically accepted way.
This seems to be to be near the opposite of what a sycophant (ie. an LLM) would do.
I've also noticed the opposite problem: Sometimes the LLM, when asked a detailed question (probably with some lead-in), pushes back in a way that betrays that they fell back to general tropes without really considering the nuances of your specific context.
This happens many times, and I usually have to lead the LLM through a chain of reasoning to prove to it that its objection, through generally sound, do not apply to my specific situation.
Someone not as well versed in the subject matter would think the LLM found a smoking gun (which they love to do), and be led on a wild goose chase.
As you say, often you check up on the LLM's "reasoning" and it doesn't follow at all, or you can easily get it to contradict itself with just as much certainty as it had about its previous convictions.
It is very scary to me that people are entrusting potentially life-altering decisions to these things.
My step mom was having debilitating pain. A year of going to doctors and no one was able to find a cause. I scanned her discharge paper work which had her prescriptions on it and gave it to Claude. It identified a prescription that had that exact side effect. They later confronted her primary care that concurred and took her off it.
A friend of mine's wife recently passed. They were chasing a suspected heart defect for over a year. She had been intermittently fainting. At about the year mark they decided to scope her digestive track. They found bleeding ulcers from cancer that was all over her body. I input her fainting symptoms into Claude and gastro impact was number two suspected after heart issues.
I have a few of other cases it's helped with. I'm not sure it could do worse than my own experience with the medical system. This is doubly true in places that lack any sort of medical care.
My mom had cancer and she was on regular, suppressive chemotherapy. I put her info into an AI and it correctly noted that her chemotherapy had stopped being effective 2 months prior based on factual lab reports. She was unaware of this. I was able to be her health advocate much more effectively by respectfully asking her oncologist targeted questions. He was already on top of it and was addressing the issue. Our conversation was respectful and, due to my educating myself, went up another level. Ultimately, it was a positive interaction. I was satisfied that he was indeed expert at his craft, and he was satisfied that we were aware of the uncertainty of the new treatment with a risk-based understanding of the viability of success. This was a positive engagement with an expert. In parallel situations around non-health issues, I've found the ego of the expert seems to be the determinative factor in whether or not the interaction goes well.
> It’s also scary how easily you can lead each LLM to the answer you have in mind.
Scary in this context of course, but I find that it is an interesting thought for coding: it suggests that maybe, a developer who knows what they are doing will end up leading the LLM to coding something that make more sense than a developer who doesn't know and just vibe-codes blindly.
And all it takes is not blindingly accepting the first thing it spews if you suspect there's a better answer (and are in a position to evaluate that better answer).
As someone who uses Claude Code to summarize published research, you have to ground it in peer-reviewed results or it gets lost. But also, I am grounded with two degrees in the source material. So I am feeding it my views and asking if the published work agrees or disagrees with my opinions and I get fantastic results that way to the point of knowing current clinical trials and treatment regimens than most of the oncologists and which led to a great conversation with the clinical trials team. This doesn't replace people, but it augments existing expertise amazingly well.
But also, I hear so many tales of running out of tokens. I ask Claude Code to build a tool to perform a task. I review the tool and then I let it rip if I'm happy with it. As I understand things, most just ask Claude Code to do the task. That seems a bit fraught.
Anyway, you have to impose constraints IMO and ask the right questions to get the answers you need or yes Claude Code (or any other LLM) will eventually just agree with you.
Yeah a lot of focus lately on making context windows enormous and putting everything in them. (It should know every detail of your life!) But in my experience LLMs are extremely "prime-able" and also tend to hyperfixate on details.
So when asking difficult questions I tend to remove as much context as possible, rather than adding it. I don't want it to reflect my own ideas or biases back to me, I want an actually fresh perspective.
The problem is how do you know whether the answer is just the most persuasive or actually the most accurate one? It's hard to figure this out without domain knowledge.
Worse is that LLMs are trained to be persuasive by default. The "you're absolutely right..." stereotype is because these things are A/B tested on response quality and we know from studies people reliably rate vibes better then anything else - e.g. while the quality of hospital accomodations likely has some impact on patient outcomes, the view and decor of the room certainly did not fundamentally change the quality of the care provided but it is the largest determinant in how well people rate that care.
Do people here not realise that "second opinions" are a thing because humans disagree with each other when presented with the same case all the time? It's not just an LLM thing!
Why should a radiologist have to debunk AI slop? They have enough to do already. That's the same mentality that is frustrating open-source repositories with sloppy pull requests, and saying "here, sort this out for me".
Depending on the disease, even in cancer there's myeloma which may cause bone metastasis in many parts of the body with very focal lesions. Radiologists can't assess each and an every one of them, or even to find them all. So AI can definitely help in these scenarios.
I do something similar with reviewing code: I have one agent write the code and another reviews it, then they go back and forth for a bit improving the code. Seems to yield better results than one agent alone.
The difference is that in the code situation, you can run unit tests on the code, compile it, etc. Unless your LLMs are ordering diagnostics and reviewing the results, there is no further information that the LLMs have on the situation. Having a second LLM review the first is counterproductive, if the 2nd LLM is better, why not use it directly? If not, then what prevents it from sending the first on some incorrect tangent?
Also, there are multiple "correct" ways to code something, so imperfect code that solves the problem is still useful. A medical diagnosis is either correct or incorrect.
Different prompt approaches and training doctors to use LLMs can improve accuracy of LLM-assisted diagnosis. It’s pretty reasonable to hypothesize that LLM “peer review” could improve that as well.
With direct discussion, the same tendency to harmonize towards groupthink applies.
Aside from the statelessness GP mentioned, one can insert anti-conciliatory intermediation. "I saw a random claim go by, but something about it seems not quite right. What am I missing? They said: [...]." Weaponizing the bias, and orchestrating the discourse from the harness.
Run it with temperature 0 if you want to minimize randomness. Sampling from a probability distribution is not a problem by itself. The problem is when the probability distribution prioritizes wrong answers.
LLMs are well suited to my (some would say annoyingly) curious nature.
when i get an answer, and my first instinct is to ask a ton of follow-ups and "what about"s. i've learned to tamp this down with fellow humans, but with LLMs its great because most of the time the response is "you're right, something doesn't add up... let me try again". i think we eventually converge on to something reasonably true
When I ask a question outside of my domain of expertise I like to ask all of the LLMs I have access to. I also create separate sessions and ask the same question multiple ways.
It’s revealing to see how many different and contradictory answers I get, most of which are presented confidently.
The last time I ran a medical question through Claude I couldn’t even get consistent answers between sessions.
It’s also scary how easily you can lead each LLM to the answer you have in mind. When I would start asking questions about different options that other LLMs had presented, each session would drift toward that explanation.