This is an all too common failure mode (of humans). The system is working as designed, being a next most probable token generator.
What I keep encountering is humans failing to grasp that the LLM has no world model, and no sense whatsoever of "meaning" or "truth" of anything, ever.
Because a lot of world model, truth-based, reasoning is implicitly encoded in language, frequently true things are probable next tokens.
This makes humans think "understanding" happens. It doesn't.