If you think about saying something you should not, your body does not say it. Speech and movement are voluntary, so something sits between the intention and the act, and whatever that is has a shape worth understanding. The framing worth keeping is to make a machine more human rather than making a person more machine-readable.
The open question. Which single action can be inferred dramatically better than it is today?
What would kill it
Without one narrow first use case this is a fascinating interface thesis and a decade of hardware. Name the action, the person, and why they would pay, or it stays a thesis.