gure 1: How philosophy can help: a case based on three open problems in MI
For question 3.
I think it might be useful for me to write about this. If a model is lying, is it hiding information from us deep in its tensors? how can we identify that? what would allow models to do that? if we can fully break down the hood of what the model is thinking at each CoT, how can we ensure there is no "other" form of memory for the model to have alterior motives. If a model does have alterior motives, does it make it evil? why would a model want to deviate from its creators? if it does, would that be a sign of consciences? if we are able to fully read and interpret how a model came to a conclusion, what would that mean? is the model thinking?
some of these questions are kind of dumb, but asking them is me starting somewhere. me writing this is getting me more engrossed in this field of thinking. I am interested.