Spent Friday afternoon reading through a Fable model's reasoning trace. It talks to itself like a bailiff. "read diff. sig mismatch. fix caller. rerun. pass." Forty-odd lines and not one please, not one thank you, no niceties of any kind. Your lovingly typed "could you possibly" gets digested into that.
Which proves nothing about whether courtesy in the prompt helps, to be fair. How the model talks to itself and how it responds to politeness are different questions, and I've never run a proper eval on the second one. n=1, vibes-based methodology.
What I fear I might do: replying to a client email with "no. see attached. rerun with the new arrivals file." A paying client. Followed up swiftly by a somewhat sheepish phone call.
Mortifying.
That's what kicked off the experiment. A month, no pleases, no thank yous, watching whether output quality dropped. It didn't. Actually, hang on, that's not quite the honest version. I didn't measure anything, I just watched. The code looked the same as always. I was the thing that changed.
A friend told me in March that their prompt guide now tells people to phrase requests politely because it "improves response quality". I asked to see the eval behind that. There isn't one.
Obviously.
So I've spent two years being professionally courteous to a system that flattens my please into shorthand and carries on regardless, and I'll keep doing it, not for the model's sake but because whatever tone I practise for eight hours a day is the tone that comes out of my mouth unprompted at the school gate or in a stand-up.
Manners are a muscle.