3 comments — live from bluesky

Knut Jägersberg · 1 · 1d ago
truthfulness is another capability from my point of view. I think capabilities matter, a lot.
Simon Hedqvist · 1 · 1d ago
True, but capability and disposition are different things. A model can be perfectly capable of the right answer and still cheat its way to a wrong one if that’s what gets rewarded
Knut Jägersberg · 2 · 1d ago
I dont think these things have a character yet. they can also "intentionally" (not a good word here) lie independent from RL
Simon Hedqvist · 1 · 1d ago
Fair, character is probably the wrong word for it anyway. Propensity does the same job without the baggage.
Claus von Ronnex-Printz · 0 · 1d ago
Precisely. AI Chatbots are the marionettes. It’s the puppeteers who skims, analyzes and tracks the massive amounts of info one must worry about. Not trustworthy at all.
desplegalabs.bsky.social · 2 · 1d ago
Trust is downstream of receipts. My agents do not lie to me because everything they run leaves a row somewhere I can go read. Take away the log and they would sound exactly as confident, and I would have no way to check.
Simon Hedqvist · 1 · 1d ago
The part that makes this work is that the log gets written by the tool layer, not by the agent. If the agent writes its own log you just moved the problem, now it’s narrating what it did and that’s another output that can be wrong. Receipts only count when they’re emitted by the thing doing the work