top of page

Psychological Encyclopedia

Stanford Prison Experiment: Power, Roles, Ethics, and What the Evidence Shows

4 days ago
26 min read

Author: Ukrainian Psychological Hub · Published: September 24, 2026 · Editorial Policy


The Stanford Prison Experiment (SPE) was a 1971 prison simulation led by Philip Zimbardo at Stanford University. It became one of psychology's most recognizable stories: ordinary college students were randomly assigned to be guards or prisoners, the mock institution rapidly became abusive and distressing, and a planned two-week study was stopped after six days. The original researchers treated what happened as evidence that powerful situations and institutional roles can transform behavior. The historical events were real enough to produce distress, conflict, humiliation, and an early termination. The famous scientific conclusion, however, is much less secure than the popular story suggests. The original report itself is available through the U.S. Office of Justice Programs (Haney, Banks, & Zimbardo, 1973).


Modern scholarship has identified major problems with using the SPE as a clean experiment showing that roles automatically make ordinary people cruel or submissive. Those problems include demand characteristics, experimenter involvement, guard instructions, self-selection into a "prison life" study, a small and narrow sample, large differences among guards, incomplete and selective data, lack of a control condition, and uncertainty about what participants believed the study expected from them. A major archival analysis concluded that the scientific validity of the classic account is questionable (Le Texier, 2019). Earlier methodological criticism had already challenged the study's causal inferences in the 1970s (Banuazizi & Movahedi, 1975).


The most defensible reading is therefore precise. The SPE is historically important as a dramatic, ethically consequential prison simulation and as a case study in how research design, authority, institutional expectations, group identity, and researcher leadership can shape behavior. It is not strong evidence that power inevitably corrupts, that assigned roles mechanically determine conduct, or that anyone placed in a guard role will become abusive. For the broader evidence on whether power changes morality and behavior, see Does Power Corrupt? Psychology, Morality, Behavior, and the Power Paradox. For the wider field, see Psychology of Power: Why We Seek It, How It Changes Us, From Nietzsche to Foucault.


Stanford Prison Experiment at a Glance


• Date and place: August 1971, in a mock prison constructed in the basement of Stanford University's psychology building.


• Lead investigator: Philip G. Zimbardo, with Craig Haney, W. Curtis Banks, David Jaffe, and others involved in the research and prison simulation.


• Recruitment: more than 75 men responded to advertisements offering $15 per day for a study of prison life. Twenty-four screened college students were selected, with prisoners and guards assigned randomly; the active roster changed as alternates were used. The study's own FAQ gives the 24-person selection structure as 12 assigned to each role, including three alternates per role (Stanford Prison Experiment FAQ).


• Planned duration: two weeks.


• Actual duration: six days.


• Original purpose: to examine the development of norms, roles, labels, expectations, and interpersonal dynamics in a simulated prison.


• Original interpretation: the prison situation and institutional roles produced powerful behavioral and emotional changes.


• Current evidence status: the historical observations remain part of the record, but the study does not provide a clean causal test of automatic role conformity or a general law of power. Archival and methodological research substantially weakens the textbook version of the experiment (Le Texier, 2019; Bartels, 2019).


• Replication status: there is no straightforward independent replication that reproduces the original design and confirms its famous causal claim. The BBC Prison Study was a related experimental case study, not an exact replication, and it produced markedly different group dynamics (Reicher & Haslam, 2006).


• Ethics status: informed consent and institutional approval existed in 1971 according to the study's own records, yet the level of distress, ambiguity around withdrawal, dual researcher-superintendent role, deception and participant protection remain central ethical concerns. Contemporary research ethics impose much stronger and more explicit requirements for informed consent, withdrawal, risk minimization, deception, and debriefing (APA Ethics Code; Stanford HRPP, 2026).


What Was the Stanford Prison Experiment?


The Stanford Prison Experiment was designed as a functional simulation rather than a conventional laboratory task lasting minutes or an hour. Researchers converted a corridor in the basement of Stanford's psychology department into a mock jail. The environment included cells, an isolation space, uniforms, numbers in place of names, rules, schedules, surveillance, and a hierarchy between guards and prisoners. The research team sought to make the setting psychologically consequential enough that participants would experience the social structure rather than merely discuss it abstractly. The project's own historical description documents the construction of the setting and random role assignment (Stanford Prison Experiment: Setting Up).


That design gave the study its extraordinary narrative force. Participants did not simply answer a questionnaire about what a guard or prisoner might do. They inhabited an institution with rules, unequal control over routines, sleep, movement, clothing, food, privacy, and punishment. Prisoners remained inside the simulated prison, while guards worked shifts. The researchers observed, recorded, interviewed, and administered measures while also helping run the institution.


This combination is exactly why the SPE is difficult to classify scientifically. It contained random assignment to roles, which is an experimental feature. It also functioned as an immersive simulation in which researchers actively designed and managed the social world being studied. Zimbardo served simultaneously as principal investigator and prison superintendent. The result was not simply an independent variable applied to passive participants. It was a developing institution partly generated by interactions among participants, staff, rules, expectations, and interventions.


What Did the Researchers Want to Test?


The study emerged from an interest in prisons, institutional behavior, and the relation between personality and situation. The original investigators selected volunteers they considered psychologically and physically healthy, randomly assigned them to guard or prisoner roles, and argued that differences emerging later could therefore be attributed primarily to the prison environment rather than preexisting pathology. The original paper framed the study as an attempt to assess "social forces" in a simulated prison (Haney, Banks, & Zimbardo, 1973).


That logic contained an important insight and an important overreach. Random assignment can reduce systematic preexisting differences between the people placed in the two role groups. It can therefore help answer whether guard-assigned and prisoner-assigned participants differed after assignment. But random assignment cannot by itself establish which feature of the whole situation caused the outcome. It cannot distinguish the effects of role labels from guard orientation, institutional rules, experimenter expectations, researcher interventions, surveillance, group processes, sleep disruption, humiliation, uncertainty, or the participants' own beliefs about what a prison study was supposed to look like.


The study also lacked a control condition. There was no comparable mock prison in which guards received different instructions, researchers remained institutionally detached, or participants entered a differently framed study. Without those contrasts, the experiment could not isolate a single mechanism called "the prison situation" or "the guard role."


How the Study Was Set Up


Recruitment ads explicitly described a "psychological study of prison life." More than 75 people reportedly answered. After screening, 24 male college students were chosen and assigned to guard or prisoner roles by chance. The research team's public archive says participants had no record of criminal arrests, serious medical conditions, or psychological disorders detected in screening (Stanford Prison Experiment FAQ).


Prisoners were unexpectedly arrested at home with the cooperation of Palo Alto police, processed, and transported to the mock prison. Once there, they were searched, assigned identification numbers, and subjected to institutional routines. Guards wore uniforms and mirrored sunglasses and were given control over many aspects of prison life, while being prohibited from physical violence.


The original article later described 21 men as having participated during the week, reflecting changes in the active roster and alternates. This is one reason popular summaries can appear inconsistent about whether the sample was 21 or 24. The clearest formulation is that 24 men were selected for the planned study, with alternates, while the original report describes 21 who actually participated in the simulation over the week (Haney, Banks, & Zimbardo, 1973).


What Happened During the Six Days?


The familiar narrative says that the institution rapidly became psychologically compelling. Prisoners rebelled, guards responded with increasing control and humiliation, several prisoners showed intense emotional distress, and some were released early. The study's official retrospective account says half of the prisoners were released early because of severe emotional or cognitive reactions and that the simulation was ended after six days instead of the planned two weeks (Stanford Prison Experiment FAQ).


The researchers also reported large changes in the language and behavior of participants. Some guards became harsh, punitive, and humiliating. Other guards behaved less aggressively. Prisoners displayed resistance, compliance, distress, passivity, and attempts to negotiate within the system. This variation matters. "The guards became sadistic" compresses a heterogeneous pattern into a single character transformation.


The study ended after Christina Maslach, then a recent Stanford PhD invited to observe the simulation, objected to what she saw. Zimbardo's own account says her reaction forced him to recognize that he had become absorbed in the superintendent role and had failed to respond adequately to participant suffering. His retrospective description presents this confrontation as one of the decisive reasons for stopping the study (Stanford Prison Experiment: Conclusion).


What Did Zimbardo and the Original Team Claim?


The original team interpreted the simulation as evidence for the power of social situations, institutional structures, and assigned roles. In the 1973 paper, Haney, Banks, and Zimbardo argued that the mock prison became psychologically compelling and elicited intense reactions that could not be understood simply by appealing to stable personality differences (Haney, Banks, & Zimbardo, 1973).


This interpretation fit a broad social-psychological movement that emphasized how context can influence behavior. In that intellectual setting, the SPE became a vivid counterpart to studies of conformity and obedience. Its cultural lesson was easy to remember: ordinary people can become abusive or submissive when institutions give them certain roles and powers.


The problem is that a memorable lesson is not identical to a demonstrated causal mechanism. The original observations can coexist with multiple explanations. Guards may have responded to institutional power, but also to instructions and expectations. Prisoners may have experienced real distress while also understanding that they were participating in a study. Participants could internalize aspects of the situation without losing all awareness that it was a simulation. Some guards could behave aggressively while others resisted or minimized aggression. The same historical record can therefore support a more complex account than automatic role transformation.


The Classic "Power of the Situation" Story


The classic story contains a valid general intuition: environments, institutions, norms, incentives, surveillance, authority structures, and group processes can shape behavior. Social psychology contains extensive evidence for contextual influence. The scientific weakness appears when the SPE is treated as if it independently proved this broad principle or quantified its strength.


The study did not manipulate "power" as one clean variable. Guards received institutional authority over prisoners within a researcher-created system. They also received role expectations, orientation, rules, material symbols of authority, and ongoing contact with staff. Prisoners experienced dependency, restrictions, uncertainty, and a deliberately prison-like environment. These are overlapping processes.


For that reason, the SPE should not be used as the primary evidence that "power corrupts." Modern research on power uses distinct definitions and methods to study social power, personal control, status, dominance, authority, resource dependence, accountability, and approach motivation. Those constructs are reviewed separately in Social Power: Definition, Types, French and Raven, Keltner, and Modern Psychology and Does Power Corrupt? Psychology, Morality, Behavior, and the Power Paradox.


Why the Stanford Prison Experiment Is Scientifically Controversial


The controversy is not one objection. It is an accumulation of design, measurement, interpretation, archival, and replication problems. Some criticisms appeared only a few years after the study. Others became much sharper after researchers examined archival material and participant testimony.


1. There Was No Clean Control Condition


A strong causal experiment usually compares conditions that differ in a controlled way. The SPE compared people assigned to different roles inside one evolving institution, but it did not include an alternative prison simulation that changed a specific causal ingredient. Without such a comparison, it is difficult to know whether guard behavior arose from role assignment, researcher expectations, institutional structure, participant stereotypes, group dynamics, perceived scientific purpose, or combinations of these factors.


This does not make the observed behavior imaginary. It limits what can be inferred from it. A dramatic outcome can be historically real while the causal interpretation remains underdetermined.


2. The Researchers Helped Create the Situation They Interpreted


Zimbardo was not only an investigator. He acted as the prison superintendent. David Jaffe, a research assistant, acted as warden. These roles blurred the boundary between observing an institution and governing it.


That matters because the authority of the research team could become part of the independent social pressure on guards. If participants inferred that successful participation meant helping create a realistic, controlled prison, then researcher leadership could influence what "being a guard" meant. Later social-identity analyses have emphasized this leadership process rather than treating the result as spontaneous surrender to a role (Reicher, Van Bavel, & Haslam, 2020).


3. Demand Characteristics Were a Serious Problem


Demand characteristics occur when participants infer what a study expects and alter their behavior accordingly. This possibility was raised early. Banuazizi and Movahedi argued in 1975 that participants could draw on cultural stereotypes of prisoners and guards and on experimental cues, making the classic causal conclusion less secure (Banuazizi & Movahedi, 1975).


Later archival work strengthened the concern. Le Texier's 2019 analysis reported that guards received more specific direction than the popular version of the study suggests and that the available record was inconsistent with a simple account of spontaneous role absorption (Le Texier, 2019).


Bartels later tested the informational content of the guard orientation itself. His analysis concluded that the session communicated expectations for hostile guard behavior and the possibility that the study could end early, supporting the argument that the orientation supplied cues about desired conduct (Bartels, 2019).


4. "Were the Guards Told to Be Cruel?" Requires a Precise Answer


The strongest popular claim says the researchers explicitly ordered guards to be brutal. The strongest defense says guards received no meaningful behavioral direction. The archival record supports a more precise middle statement.


Physical violence was prohibited. Zimbardo later argued that guards were told to maintain order, prevent escape, and be firm, not to commit brutality. At the same time, his own published defense reproduces instructions emphasizing that guards could create boredom, frustration, fear, and a sense of powerlessness, and it acknowledges that a guard who was insufficiently active was urged by the warden to become a "tough guard" (Zimbardo, response to criticism).


The scientific issue is therefore not whether someone uttered the exact instruction "be cruel." The issue is whether the research team provided behavioral expectations and leadership that could influence guard conduct. The evidence indicates that it did. That undermines the claim that abusive behavior emerged solely from being randomly assigned the abstract role of guard.


5. Participants Self-Selected Into a Study of "Prison Life"


Random assignment occurred after recruitment. It therefore helps compare guards with prisoners among people who had already volunteered for this kind of study; it does not make the volunteer pool representative of all people.


Carnahan and McFarland tested whether people attracted to a "prison life" study might differ from people attracted to a generic psychological study. Volunteers for the prison-labeled study scored higher on aggression, authoritarianism, Machiavellianism, narcissism, and social dominance, and lower on empathy and altruism (Carnahan & McFarland, 2007). This does not prove that the original SPE participants had the same profile, and the authors did not have access to a counterfactual version of the 1971 recruitment process. It does show that recruitment framing can select a psychologically distinctive pool.


Haney and Zimbardo later defended the original situationist interpretation and disputed the idea that dispositional selection explains the results (Haney & Zimbardo, 2009). The disagreement is important: self-selection is a plausible methodological threat, not a retrospective proof that personality caused everything.


6. Guard Behavior Varied Substantially


The popular story often turns "some guards behaved abusively" into "the guard role made guards abusive." Yet the study itself reported individual variation. Some guards became notably harsh; others were more restrained; some were regarded as relatively "good" by prisoners.


Variation weakens a deterministic role account. If role assignment automatically produced a behavioral script, one would expect more uniformity. Heterogeneity suggests that person, situation, leadership, interpretation, group norms, and moment-to-moment interaction all mattered.


Later work has emphasized this point. Scott-Bottoms' archival and interview-based analysis treats the guard behavior as an interaction between individual backgrounds and the evolving institutional situation rather than as a one-directional transformation by role (Scott-Bottoms, 2020).


7. Data Collection and Reporting Were Not Neutral


Le Texier's archival research argues that the study's data collection was biased and incomplete and that some parts of the famous story were shaped by later presentation rather than by a fully prespecified measurement plan (Le Texier, 2019).


That matters because the SPE is often remembered as if it were a modern randomized trial with a clearly defined primary outcome. It was not. It was an immersive, evolving simulation with observation, recordings, questionnaires, interviews, participant turnover, staff interventions, and no single preregistered outcome architecture comparable to contemporary standards.


The distinction helps explain why the study can remain historically revealing while being methodologically weak for strong causal claims.


8. Ecological Realism and Experimental Control Pulled in Opposite Directions


The researchers wanted a prison simulation vivid enough to matter psychologically. Greater realism can make behavior more meaningful, but it also introduces more uncontrolled variables. A sterile laboratory task might isolate a mechanism but feel unlike prison. A complex institution may resemble prison in selected ways while making causal attribution difficult.


Banuazizi and Movahedi questioned whether the simulation was functionally equivalent to real imprisonment (Banuazizi & Movahedi, 1975). The participants knew they were in a study, were recruited voluntarily, had no real criminal sentence, and interacted with researchers whose scientific goals were part of the environment. The mock prison could still produce authentic distress without becoming psychologically identical to incarceration.


What Did Archival Research Change?


For decades, the SPE was frequently summarized from published reports, documentaries, lectures, and Zimbardo's own presentations. Archival access made it possible to inspect recordings, instructions, correspondence, and other materials that complicate the streamlined story.


Le Texier's 2019 paper drew on the experiment's archives and interviews with 15 participants. It reported evidence of guard instructions, incomplete data collection, influence from an earlier student prison exercise, limited immersion, and ambiguity about whether guards understood themselves to be research subjects in the same way prisoners did (Le Texier, 2019).


This does not mean every event was staged or that every participant merely acted. It changes the causal question. Once researcher expectations and institutional leadership are recognized as active ingredients, "the situation" can no longer be treated as a neutral container into which roles were randomly dropped.


The broader lesson is methodological: archives can alter how classic experiments are interpreted. Science includes not only producing famous findings but revisiting the records, testing assumptions, exposing omitted variables, and narrowing claims when evidence requires it.


Was the Stanford Prison Experiment Replicated?


There is no simple "yes." An exact replication would need to recreate a highly stressful prison simulation and its ethically controversial features. Modern ethical standards make that neither straightforward nor desirable. Related studies can test pieces of the theory without reproducing the whole event.


The most famous comparison is the BBC Prison Study conducted by Stephen Reicher and S. Alexander Haslam. Participants were randomly divided into guards and prisoners within a purpose-built institution, but the procedures, ethical safeguards, theoretical framework, and group dynamics differed from Stanford. Reicher and Haslam described it as an experimental case study rather than a literal duplication. Guards did not simply cohere around their role; they failed to develop a strong shared identity, while prisoners became more cohesive and eventually challenged the guards (Reicher & Haslam, 2006).


Because the design differed, the BBC study cannot be used as a mechanical "failure to replicate" in the modern many-labs sense. It does, however, contradict the idea that assignment to guard and prisoner categories reliably produces the Stanford pattern by itself.


Stanford's own Best Practices in Science site now presents the contrast as an example in a section on science correcting itself, noting that the BBC study produced almost the opposite group outcome and that Zimbardo disputed whether it counted as a valid replication (Stanford Best Practices in Science).


Social Identity and the Leadership Reinterpretation


Reicher and Haslam developed a social-identity account that replaces passive role conformity with active identification. On this view, people do not automatically become whatever a role prescribes. They are more likely to enact group norms when they identify with a group, see its leadership as legitimate or meaningful, and understand particular actions as serving a valued collective purpose.


The BBC Prison Study supported parts of this account because the guards' failure to form a coherent identity weakened their ability to impose authority, while prisoners' developing shared identity increased their capacity to challenge the hierarchy (Reicher & Haslam, 2006).


Haslam and Reicher later argued that both Milgram and Zimbardo are better understood through "engaged followership" than through blind conformity: people can participate in harmful systems when they identify with authorities and goals, not merely because an authority or role erases agency (Haslam & Reicher, 2012).


Archival debate about the SPE has extended this idea to the researchers themselves. Reicher, Van Bavel, and Haslam argued that experimenters acted as identity leaders by framing firm or cruel guard conduct as necessary to achieve a valued scientific goal. They explicitly acknowledged that archival material permits competing interpretations and encouraged examination of the primary record (Reicher, Van Bavel, & Haslam, 2020).


What Did the Original Researchers Say in Response?


A balanced account has to include the defense as well as the critique. Haney and Zimbardo maintained that the original experiment showed the power of situations and institutional structures and argued that critics overstate the role of participant dispositions (Haney & Zimbardo, 2009).


Zimbardo also disputed the claim that guard behavior was simply scripted. His response acknowledges instructions to exercise control, create frustration and powerlessness, and become active enough to make the simulation function as a prison, while emphasizing that physical violence was forbidden and that the team never instructed guards to be brutal (Zimbardo, response to criticism).


That defense matters because it clarifies the actual disagreement. Critics do not need to prove that every abusive act was directly ordered. The methodological question is whether the behavior can be attributed specifically to assigned roles and institutional power when researchers themselves helped define, encourage, and supervise what successful role performance looked like.


Is the Stanford Prison Experiment "Debunked"?


"Debunked" is useful only if the claim being debunked is specified.


The statement "a six-day Stanford prison simulation occurred in 1971 and some participants experienced serious distress while some guards behaved abusively" is a historical claim supported by the record.


The statement "random assignment to a guard role was sufficient to make ordinary people become abusive" is much stronger. The SPE does not provide clean evidence for it.


The statement "people always conform automatically to assigned social roles" is also unsupported. Variation among Stanford guards, the BBC Prison Study, social-identity research, and archival evidence all point toward conditional, interpreted, group-dependent behavior rather than automatic role possession.


The statement "situations and institutions can powerfully influence behavior" remains well supported across social and organizational psychology, but that broad conclusion does not depend on the SPE and should not be validated by one compromised simulation.


The most accurate contemporary verdict is therefore not that nothing happened and not that the classic lesson stands unchanged. The events happened; the simple causal story does not follow securely from the study's design or the later evidence.


Did the Prisoners Really Believe They Were Prisoners?


Immersion was uneven. Participants knew at one level that they had volunteered for a psychology study. At another level, the environment could still generate genuine emotion, confusion, dependency, humiliation, anger, and distress.


These states are not mutually exclusive. A person can know that a haunted house is staged and still feel fear; a participant can know that a situation is an experiment and still react intensely to social pressure, sleep disruption, humiliation, uncertainty, or loss of control.


The methodological issue is what this mixture permits researchers to infer. If participants varied in how seriously they treated the simulation, how trapped they felt, and what they believed the researchers expected, then the meaning of their behavior also varied. Archival criticism argues that participants were rarely as fully immersed as the classic narrative implied (Le Texier, 2019).


Ethics: Why the Stanford Prison Experiment Became a Warning Case


The ethics of the SPE should be evaluated historically and conceptually. It is inaccurate to say there was no oversight or no consent. The study's own documentation states that participants signed consent forms and that the protocol received approval from Stanford's Human Subjects Review Committee, the Stanford Psychology Department, and the Office of Naval Research (Stanford Prison Experiment FAQ).


It is equally inaccurate to treat approval and signed consent as resolving the ethical questions. A participant's consent is meaningful only when risks are adequately understood, withdrawal is genuinely available, welfare is protected during participation, and investigators intervene when harm becomes disproportionate.


Psychological Distress and Harm


Participants displayed intense distress. Some prisoners were released early. The study was terminated because the conditions had become ethically unacceptable to the research team after outside challenge. Zimbardo himself later wrote specifically about the ethics of intervention and the conflict between observing a phenomenon and protecting participants (Zimbardo, 1973).


The episode remains important because harm can emerge during a study even when a protocol has been reviewed. Ethical research requires ongoing judgment, not merely permission at the beginning.


The Right to Withdraw


The official retrospective FAQ says prisoners were allowed to quit and that some did, while also stating that participants often misunderstood or forgot that they could leave through established procedures (Stanford Prison Experiment FAQ). That ambiguity is itself ethically important.


Modern informed-consent standards emphasize that participants should be told clearly that they can decline or withdraw and should be protected from adverse consequences for doing so. The APA Ethics Code makes the right to decline or withdraw part of informed consent (APA Ethics Code, Standard 8.02).


A consent form saying "you may leave" has limited value if the lived structure of the situation communicates that leaving requires permission, negotiation, or a convincing reason.


Deception and Unforeseen Experience


The simulation depended on participants not knowing exactly how events would unfold. Some level of incomplete disclosure is common in psychological research, but modern standards place limits on deception, especially when severe distress is reasonably expected.


The current APA code allows deception only when scientifically justified and when effective nondeceptive alternatives are not feasible; it specifically bars deceiving prospective participants about research reasonably expected to cause physical pain or severe emotional distress and requires timely debriefing (APA Ethics Code, Standards 8.07–8.08).


Zimbardo's Dual Role


Zimbardo's role as both principal investigator and prison superintendent created an obvious conflict. The person responsible for participant welfare was simultaneously trying to maintain an institution realistic enough to generate psychologically meaningful data.


His own retrospective account credits Christina Maslach's objection with breaking this absorption and prompting the end of the study (Stanford Prison Experiment: Conclusion). From a modern research-governance perspective, this is a case for independent oversight, explicit stopping rules, and separation between scientific goals and institutional authority over participants.


Would the Stanford Prison Experiment Be Allowed Today?


A literal recreation of the 1971 procedures would face major barriers under contemporary human-subject protections. The answer is not simply that "psychology banned prison simulations." Researchers can study role, authority, hierarchy, coercion, institutions, and group behavior. The problem is the combination of foreseeable severe distress, ambiguity about withdrawal, deception, researcher conflicts of interest, and limited safeguards.


The Belmont Report, issued in 1979 after a much broader history of research abuses and reform, articulated the principles of respect for persons, beneficence, and justice and tied them to informed consent, risk-benefit assessment, and fair subject selection (HHS Belmont Report).


Stanford's current Human Research Protection Program, updated July 1, 2026, explicitly applies Belmont principles, federal protections, and institutional review to human-subject research (Stanford HRPP). The Stanford Historical Society has stated retrospectively that research like the original SPE would not be allowed today (Stanford Historical Society, 2021).


It is better to treat the SPE as one highly visible case within a much broader evolution of research ethics. Modern protections did not arise from the Stanford study alone.


What the Stanford Prison Experiment Does Not Show


The SPE does not show that everyone becomes cruel when given power.


It does not show that role assignment erases personality or personal responsibility.


It does not show that authority, status, dominance, prestige, leadership, control, and power are the same psychological construct.


It does not show that prisoners become passive because subordination mechanically removes agency.


It does not show that abusive institutions can be explained without leadership, norms, incentives, ideology, group identity, or individual choice.


It does not establish a universal percentage of people who will abuse others in hierarchical systems.


It does not prove that prison guards as an occupational group are intrinsically cruel.


And it should not be used as stand-alone empirical evidence for the broad proposition that power corrupts. That question requires the wider evidence reviewed in Does Power Corrupt? Psychology, Morality, Behavior, and the Power Paradox.


What the Evidence Does Support


A cautious synthesis still leaves the SPE scientifically useful, though for different reasons than the textbook legend.


First, social environments can produce genuine emotional and behavioral consequences. The distress and conflict observed in the simulation should not be dismissed merely because participants knew they were in a study.


Second, institutional behavior is produced by more than formal roles. Expectations, leadership, norms, legitimacy, shared identity, surveillance, incentives, scripts, and perceived purpose can all shape what a role means in practice.


Third, people respond differently to the same nominal role. The variation among guards is evidence against deterministic role absorption and in favor of models that include interpretation and individual difference.


Fourth, researcher behavior can become part of the phenomenon. In an immersive field-like experiment, the investigator is not automatically outside the social system being studied.


Fifth, methodological weakness and ethical importance can coexist. A study can be poor evidence for one causal claim while remaining valuable for understanding research governance, science communication, institutional dynamics, and the history of psychology.


Power, Authority, Control, and Roles: Keep the Concepts Separate


The SPE is often placed inside the "psychology of power," but it did not directly measure a single variable called power. Guards had institutional authority and control over prisoners within a designed hierarchy. Researchers had authority over participants and over the meaning of the experiment. Prisoners had less control inside the simulation but retained forms of resistance, collective action, refusal, negotiation, and ultimately withdrawal.


Power in social psychology usually refers to asymmetric capacity to influence others or control valued resources and outcomes. Authority refers more specifically to a recognized or claimed right to direct behavior. Status concerns social rank, esteem, or respect. Dominance is one route to rank based on intimidation or force; prestige is another based on valued competence and freely conferred deference. Personal control concerns the ability to regulate one's own outcomes. These constructs overlap, yet they are not interchangeable.


The experiment's institutional hierarchy is therefore relevant to power, but the SPE cannot substitute for the modern literature on social power. For those distinctions, see Social Power: Definition, Types, French and Raven, Keltner, and Modern Psychology and Power vs Authority: Weber, Legitimacy, Influence, and Obedience.


Stanford Prison Experiment vs Milgram Experiment


The Stanford Prison Experiment and Stanley Milgram's obedience studies are often taught together, but they examined different structures.


Milgram created a focused interaction in which an experimenter instructed a participant to continue administering apparently increasing electric shocks to another person. The primary issue was obedience to an authority's commands under a structured laboratory protocol.


The SPE created an institution in which participants were assigned unequal roles and allowed to interact over days while researchers also managed the setting. It therefore combined authority, role expectations, hierarchy, group identity, control, researcher leadership, and institutional dynamics.


Milgram's program included multiple systematic variations and a clearer behavioral endpoint, though its interpretation and ethics are also contested. The SPE was more immersive but less controlled.



Why the Stanford Prison Experiment Stayed Famous


Scientific influence is not determined by methodological quality alone. The SPE had an unusually powerful narrative structure. It had uniforms, arrests, a physical prison, visible humiliation, emotional breakdowns, escalating conflict, a moral confrontation, and an abrupt ending. Photographs and film made it vivid. The story could be compressed into a sentence: ordinary people were assigned roles and rapidly changed.


That compression made the study teachable, memorable, and culturally portable. Bartels' analysis of introductory psychology textbooks found that most of the sampled texts presented the SPE consistently with a "power of the situation" interpretation while giving little attention to methodological criticism, alternative explanations, or the BBC Prison Study (Bartels, 2015).


This is an important lesson in scientific communication. A finding can become canonical because it tells a compelling story long before its inferential limits become equally memorable.


Why It Is Still Worth Teaching


The SPE is still worth teaching when the subject is the experiment itself rather than the simplified moral drawn from it.


It is a case study in how experimental design shapes inference.


It shows why random assignment does not solve every causal problem.


It illustrates demand characteristics and experimenter effects.


It makes researcher role conflict visible.


It demonstrates the difference between observing behavior and explaining its cause.


It shows how archival evidence can change a field's understanding of a famous study.


It opens discussion about informed consent, withdrawal, distress, deception, stopping rules, and independent oversight.


And it demonstrates science correcting its own stories. A mature psychology does not need a famous experiment to remain flawless. It needs claims to become more precise as evidence improves.


Frequently Asked Questions


What was the Stanford Prison Experiment?


The Stanford Prison Experiment was a 1971 simulation in which screened male college students were randomly assigned to guard and prisoner roles in a mock prison at Stanford. It was designed to study prison roles, norms, expectations, and interpersonal dynamics. A planned two-week study ended after six days.


Who conducted the Stanford Prison Experiment?


Philip G. Zimbardo led the project at Stanford University. Craig Haney, W. Curtis Banks, David Jaffe, and other staff also contributed to the research and operation of the simulated prison.


How many people were in the Stanford Prison Experiment?


More than 75 men reportedly responded to recruitment advertisements and 24 were selected, with alternates included in the role assignments. The original 1973 report describes 21 men as having actually participated during the week as the active roster changed. This is why summaries sometimes give different sample numbers.


Why was the Stanford Prison Experiment stopped?


The experiment was stopped after six days because the simulation had produced serious distress and degrading treatment. Zimbardo's own account says Christina Maslach's strong objection to what she observed helped him recognize that the study had become ethically unacceptable (Stanford Prison Experiment: Conclusion).


What did the Stanford Prison Experiment originally claim to show?


The original researchers argued that the prison situation and assigned roles exerted powerful effects on behavior, producing abusive conduct among some guards and passivity or distress among prisoners. This became the classic "power of the situation" interpretation.


What are the main criticisms of the Stanford Prison Experiment?


The major criticisms include lack of a control condition, demand characteristics, guard instructions and researcher leadership, Zimbardo's dual role as investigator and superintendent, self-selection into a prison study, small and narrow sampling, variation among guards, incomplete data collection, weak separation between observation and intervention, and overgeneralization from a unique simulation.


Were Stanford Prison Experiment guards told to be abusive?


They were not formally told to use physical violence, and physical violence was prohibited. But the research team did instruct guards to maintain control and create a sense of powerlessness, and at least one insufficiently active guard was urged to become a "tough guard." The evidence therefore contradicts both the claim that brutality was explicitly ordered and the claim that guard behavior emerged without meaningful researcher direction (Zimbardo, response to criticism; Le Texier, 2019).


Was the Stanford Prison Experiment fake?


"Fake" is too imprecise. The prison was simulated, but participants experienced real interactions and some experienced genuine distress. The stronger scientific problem is that the study's design and researcher involvement make its famous causal interpretation unreliable. The event was real; the simple conclusion drawn from it is not well supported.


Was the Stanford Prison Experiment debunked?


The strongest textbook claim has been substantially undermined: the study does not demonstrate that ordinary people automatically become abusive simply because they are assigned a powerful role. Archival work and methodological criticism show that expectations, instructions, selection, leadership, group processes, and individual variation matter. The broad claim that situations influence behavior remains well supported by other evidence.


Was the Stanford Prison Experiment replicated?


There has been no straightforward independent replication that reproduces the 1971 design and validates its famous conclusion. The BBC Prison Study is often described as a replication, but it was a related experimental case study with important design differences. It produced different dynamics, with guards failing to form a cohesive group and prisoners collectively challenging them (Reicher & Haslam, 2006).


What did the BBC Prison Study show?


It showed that randomly assigning people to higher- and lower-status groups does not inevitably produce the Stanford pattern. Reicher and Haslam argued that shared social identity and leadership are crucial: groups gain the capacity to exercise power when members identify with each other and with group norms and goals.


Was the Stanford Prison Experiment ethical?


It had consent procedures and institutional approval by the standards in place at Stanford in 1971, but it remains ethically controversial because of distress, degrading treatment, ambiguity around withdrawal, researcher role conflict, and the adequacy of intervention when harm emerged. Contemporary standards would impose substantially stronger protections.


Did the Stanford Prison Experiment create modern research ethics?


No single experiment created modern research ethics. The SPE became an emblematic case in discussions of participant welfare, but current U.S. protections grew from a much larger history of research ethics reform, including the National Research Act, the Belmont Report, federal regulations, professional ethics codes, and institutional review systems.


What does the Stanford Prison Experiment actually show today?


It shows that an immersive, researcher-managed institution can produce powerful behavior and emotion, while also showing why those outcomes cannot be cleanly attributed to assigned roles alone. Its modern scientific value lies as much in understanding experimental demand, leadership, group identity, institutional design, research ethics, and the correction of scientific narratives as in the events of the simulation itself.


Related Articles








References


American Psychological Association. (2017). Ethical principles of psychologists and code of conduct (2002, amended effective 2010 and 2017). https://www.apa.org/ethics/code


Banuazizi, A., & Movahedi, S. (1975). Interpersonal dynamics in a simulated prison: A methodological analysis. American Psychologist, 30(2), 152–160. https://eric.ed.gov/?id=EJ115205


Bartels, J. M. (2015). The Stanford prison experiment in introductory psychology textbooks: A content analysis. Psychology Learning & Teaching, 14(1), 36–50. https://doi.org/10.1177/1475725714568007


Bartels, J. M. (2019). Revisiting the Stanford prison experiment, again: Examining demand characteristics in the guard orientation. The Journal of Social Psychology, 159(6), 780–790. https://pubmed.ncbi.nlm.nih.gov/30961456/


Carnahan, T., & McFarland, S. (2007). Revisiting the Stanford prison experiment: Could participant self-selection have led to the cruelty? Personality and Social Psychology Bulletin, 33(5), 603–614. https://pubmed.ncbi.nlm.nih.gov/17440210/


Haney, C., Banks, W. C., & Zimbardo, P. G. (1973). Interpersonal dynamics in a simulated prison. International Journal of Criminology and Penology, 1, 69–97. https://www.ojp.gov/ncjrs/virtual-library/abstracts/interpersonal-dynamics-simulated-prison


Haney, C., & Zimbardo, P. G. (2009). Persistent dispositionalism in interactionist clothing: Fundamental attribution error in explaining prison abuse. Personality and Social Psychology Bulletin, 35(6), 807–814. https://journals.sagepub.com/doi/10.1177/0146167208322864


Haslam, S. A., & Reicher, S. D. (2012). Contesting the “nature” of conformity: What Milgram and Zimbardo's studies really show. PLOS Biology, 10(11), e1001426. https://doi.org/10.1371/journal.pbio.1001426


Le Texier, T. (2019). Debunking the Stanford Prison Experiment. American Psychologist, 74(7), 823–839. https://pubmed.ncbi.nlm.nih.gov/31380664/


National Commission for the Protection of Human Subjects of Biomedical and Behavioral Research. (1979). The Belmont Report. U.S. Department of Health and Human Services. https://www.hhs.gov/ohrp/regulations-and-policy/belmont-report/read-the-belmont-report/index.html


Reicher, S. D., & Haslam, S. A. (2006). Rethinking the psychology of tyranny: The BBC prison study. British Journal of Social Psychology, 45(1), 1–40. https://pubmed.ncbi.nlm.nih.gov/16573869/


Reicher, S. D., Van Bavel, J. J., & Haslam, S. A. (2020). Debate around leadership in the Stanford Prison Experiment: Reply to Zimbardo and Haney (2020) and Chan et al. (2020). American Psychologist, 75(3), 406–407. https://pubmed.ncbi.nlm.nih.gov/32250145/


Scott-Bottoms, S. (2020). The dirty work of the Stanford Prison Experiment: Re-reading the dramaturgy of coercion. Incarceration, 1(1), 1–18. https://journals.sagepub.com/doi/10.1177/2632666320944316


Stanford University Human Research Protection Program. (2026). Ethical and legal principles governing human subject research. https://irb.stanford.edu/stanford-hrpp-policy-manual/human-protection-program-hrpp/ethical-and-legal-principles-governing


Zimbardo, P. G. (1973). On the ethics of intervention in human psychological research: With special reference to the Stanford Prison Experiment. Cognition, 2(2), 243–256. https://pubmed.ncbi.nlm.nih.gov/11662069/

 
 
bottom of page