VK Singh Vashisht Institute for Critical Studies
A digital think tank and production house.

Mind as Policy

Reinforcement Learning, Belief, and the Old Technology of Changing Yourself

Editorial Board · Computational Psychology · · 2,194 words · ICS-2026-154

Man tu jot saroop hai — apnaa mool pachhaan. “O mind, you are the embodiment of light — recognize your origin.” — Guru Amar Das, Guru Granth Sahib

A folk conviction and a technical field appear, at first, to flatly contradict each other, and the contradiction is worth taking seriously because both sides are right. The folk conviction is that beliefs drive action: change what a person holds to be true and good, and their behavior follows. It is the premise of every scripture, every therapy, every argument ever made in the hope of changing what someone does by changing what they think. The technical field is reinforcement learning, which models a mind as a policy — a mapping from situations to actions — shaped not by belief at all but by reward: do the thing, get the signal, update the tendency, repeat. On the reinforcement-learning account, action does not follow from belief. Belief, if it exists in the model at all, is a late rationalization of a policy that reward already wrote. This essay argues that the contradiction is an illusion of vantage point, and that the traditions which train the mind have been doing reward engineering for three thousand years without the vocabulary.

I. The Wager

The wager is that “mind as policy” is a genuine lens on the psyche — powerful, generative, and worth pushing hard — while never being the ground truth it can pretend to be. Reinforcement learning offers a substrate that recurs across the mind’s operations: inspection (attending to some states and not others), checking (evaluating outcomes against expectation), reward design (what the system has been trained to want), and simulation (rehearsing action before taking it). The claim is that this substrate reconciles the belief-first and the reward-first accounts of human change, because a policy is exactly the place where belief and reward meet — and that the reconciliation is not merely academic. It tells you how change actually happens, and therefore why the disciplines built to change people take the shape they do. The discipline the essay imposes on itself is the “not ground truth” clause: the mind is modeled well by a policy, which is not the same as being one, and the essay’s final section is about the difference.

II. What a Policy Is

Strip reinforcement learning to its frame. An agent is in a state; it takes an action; the world returns a reward and a new state; the agent adjusts, over many such cycles, toward actions that earn more reward over time. The learned object is the policy — the agent’s disposition, its answer to “in a situation like this, what do I do?” A policy is not a belief and not a rule. It is a tuned tendency, built by consequence rather than by argument, and it can be extraordinarily sophisticated without the agent being able to state a single principle it follows.

This is already a description of most of a human life. The overwhelming majority of what a person does in a day is policy in exactly this sense: driving, speaking a first language, reacting to a tone of voice, reaching for the phone. None of it is deliberated; all of it was trained by consequence; and the person could not, if asked, produce the rules. The reinforcement-learning lens earns its keep immediately by naming the largest and least-examined part of the self — not the beliefs we can recite but the dispositions we cannot, the policy that runs the animal while the reasoning mind narrates.

III. The Reversal

Here is where the folk conviction seems to lose. If the policy is written by reward and not by belief, then “change your mind and your behavior will follow” has the arrow backwards. What changes behavior is reward and repetition — new consequences, practiced until the tendency shifts. Belief looks, from inside the model, like a spectator: the story the narrating mind tells about a policy it did not author and cannot directly edit. This is the uncomfortable finding that behaviorism reached a century ago and that reinforcement learning has since formalized: you do not think your way out of a trained disposition. You cannot argue a phobia away, or reason yourself calm, or believe your way past a craving, because the thing in charge is not listening to arguments. It is a policy, and a policy is changed the way it was made — by consequence, over time.

And yet the folk conviction refuses to die, because it keeps being right in practice. People do change what they do by changing what they hold true — converts, the newly disciplined, the person who reads one book and reorders a life. Any honest model has to account for the fact that belief-change sometimes moves behavior when nothing was reinforced at all. The reinforcement-learning account, taken as ground truth, cannot. Which is the first sign that it is a lens and not the floor.

IV. Where Belief Re-enters

Belief re-enters through the parts of the frame that behaviorism left out and reinforcement learning had to put back: the model and the reward function themselves.

A purely model-free agent learns only by direct trial — it must actually take the action and feel the consequence. But sophisticated agents are model-based: they carry an internal simulation of the world and can rehearse actions inside it, earning imagined reward and updating the policy without ever moving. This is the technical home of belief. What a person holds to be true is their world-model, and the world-model is what simulation runs on. Change the model — persuade someone that an action they never tried will lead somewhere they want to go — and you have changed the input to every future simulation, and therefore the policy those simulations shape. Belief moves behavior not by overriding the policy directly but by rewriting the world inside which the policy is rehearsed. The convert’s life reorders because the map did.

And the reward function is not fixed. What a person finds rewarding — status, or peace, or the approval of a particular face, or the absence of craving — is itself trainable, and a belief about what is worth wanting is a claim on the reward function. To come to believe that the approval of strangers is worthless is not a spectator’s idle opinion; it is a re-weighting of the signal that trains every subsequent policy. This is why the reversal is only half true. Reward writes the policy, yes — but belief writes the model the policy is rehearsed against and bids on the reward the policy is trained toward. The arrow runs both ways, through two different doors.

V. Mind-Training as Reward Engineering

Now the traditions come into focus, and they were never naïve. The great mind-training systems — the dharmic disciplines, the contemplative practices, Gurmat’s insistence on simran and sangat and the daily nitnem — do not primarily work by argument, and they know it. They work by practice: repetition, in a structured environment, over long time, with the explicit aim of retraining what the self finds rewarding and reshaping the world-model it simulates against. This is reward engineering, described from the inside.

Consider the components. The disciplines install a practice (repeated action, the only thing that moves a policy). They embed the practitioner in a sangat, a community — because the reward signal that trains a social animal is overwhelmingly social, and a company of the like-minded is a reward function you can walk into. They supply a doctrine — a world-model, a picture of what is real and what is worth wanting, which becomes the substrate every simulation runs on. And they aim, explicitly, at the reward function itself: the whole project of loosening attachment is a project of retraining what the system craves, so that the policy trained downstream is different in kind. “Change your beliefs, change your animal nature” is not a pious slogan when read this way. It is an accurate, if compressed, description of a two-channel intervention: rewrite the model so simulation points elsewhere, and retrain the reward so the policy is built toward a different good. The traditions reached the architecture of reinforcement learning by the only route available before the mathematics — by trying, for centuries, to actually change people, and keeping what worked.

This also explains a fact the argument-first view of religion cannot: why the traditions are so heavy on repetition and so light, relatively, on persuasion. If belief alone moved behavior, a single convincing sermon would suffice. It never does. The traditions build cathedrals of repetition — daily prayer, recitation, the returning company — because they understood, without the vocabulary, that a policy is changed the way it was made. They are not asking you to be persuaded. They are asking you to practice, which is the only thing the animal responds to.

VI. The Not-Ground-Truth Clause

The lens has to be held at arm’s length precisely because it is so good. That the mind is modeled well by a policy tuned on reward does not make the mind a reinforcement-learning agent, and three cautions keep the model honest. First, reward in a machine is a scalar the designer specifies; reward in a person is not given from outside but is itself part of what is in question, contested, and revisable — a human can come to disown the very thing that rewards them, which is a move no clean reinforcement-learning agent makes. Second, the epigraph’s claim — that the mind is jot saroop, of the form of light, with an origin to recognize — is a claim the model cannot evaluate and should not pretend to refute; a lens that captures the policy says nothing about whether there is a perceiver the policy serves. Third, and practically: to treat the model as ground truth is to treat people as trainable objects, and the same mathematics that illuminates self-discipline becomes, in other hands, the engineering of the feed, the slot machine, the manufactured craving. The lens that lets you retrain yourself toward freedom is the lens that lets a platform retrain you toward the scroll. What decides which it becomes is not in the mathematics. It is in who holds the reward function, and to what end.

Counter-case

The strongest objection is that this reconciliation is too neat — that it “explains” the belief-first and reward-first accounts by relabeling belief as “model” and “reward-weighting,” and that relabeling is not reconciling. A behaviorist would say the model-based machinery is unnecessary epicycles and that everything reduces to conditioning in the end; a cognitivist would say the reduction to policy loses exactly what matters about a reasoning mind. Both would say the essay wants to have it both ways.

The essay does want both ways, and defends the wanting, because the phenomenon has both ways in it. The pure behaviorist cannot explain the convert whose behavior reorders with no reinforcement; the pure cognitivist cannot explain the phobia that reasoning cannot touch. A model that captures only one is falsified by the other. The value of “mind as policy” is precisely that it locates where each account is true — reward writes the disposition, belief writes the model and bids on the reward — rather than declaring one the whole story. That is not having it both ways as an evasion. It is having it both ways because the mind is a loop, and a loop, cut at either point, tells a false story about a true circle.

Stakes

Computational Psychology, as the institute practices it, is the wager that the vocabularies of computation can illuminate the psyche without colonizing it — that “policy,” “reward,” and “model” can be genuine instruments rather than a new reductionism wearing mathematics. This essay is a test of that wager on the oldest question the psyche poses: can a person change, and if so, how. The reinforcement-learning lens answers with unusual precision — change is the slow rewriting of a policy through practice, model, and reward — and in answering, it vindicates rather than dissolves the mind-training traditions, which turn out to have been engineering exactly these variables all along. The danger and the promise are the same fact: the mind is, to a real approximation, trainable. The traditions used that fact to make people free. The feed uses it to make them stay. The mathematics is indifferent between them, which is why the question the model cannot answer — trained toward what, and by whom — is the only one that finally matters.


Sources

  • The reinforcement-learning frame: agent, state, action, reward, policy; model-free versus model-based learning (standard formulations in the field).
  • The behaviorist tradition on conditioning and the limits of persuasion; the cognitivist reply on world-models and simulation.
  • On mind-training as practice, community (sangat), doctrine, and the retraining of desire: Gurmat tradition (simran, nitnem, sangat); dharmic contemplative disciplines. The epigraph: man tu jot saroop hai (Guru Amar Das, Guru Granth Sahib).
  • On the industrial capture of the reward function: the institute’s Compound Is Your Couch.

Suggested citation

Editorial Board. “Mind as Policy.” VK Singh Vashisht Institute for Critical Studies, September 2026. vashisht.institute/essays/mind-as-policy. ICS-2026-154.