What is a criminal? What is a crime? What purpose do these categories serve? These are the questions I can't help but ask as Claude's watermarking technology makes headlines. At first, they seem totally unrelated to the topic of AI disclosure. After all, we are merely dealing with issues of authorship and attribution; what does lawlessness have to do with it? Quite a lot, perhaps, though I did not arrive at that conclusion directly.
I had been thinking about AI disclosure for a different essay, specifically about people who genuinely dislike reading AI-generated content. I might disagree with some of the assumptions behind their dislike, but their preferences are nonetheless real. If knowing that AI produced or substantially assisted with a piece of writing would change someone's willingness to read it, there is an argument to be made that upfront disclosure honors that preference. The thought seemed straightforward enough until I considered what disclosure actually asks of the person doing the disclosing.
Disclosure does not occur in a vacuum. "AI was used to produce this" is not received as a neutral statement of provenance in a culture where AI use can already mean lazy, cheater, hack, grifter, cognitively deficient, environmentally irresponsible, or incapable of writing and thinking independently. Disclosure makes a person legible to a category, and categories come with consequences. It is, therefore, unsurprising that some people might prefer to hide their AI use, particularly when they believe that what they actually did with the technology bears little resemblance to the caricature awaiting them on the other side of disclosure.
My first reaction to this possibility was almost embarrassingly conventional: if you do something bad, you have to be willing to bear the consequences of the act. Then I stopped. There was an enormous assumption hiding inside that sentence. Before someone can deserve the consequences of doing something bad, we have to have correctly established that the thing was bad, that the consequences correspond to the wrong, and that whoever is imposing them possesses some legitimate authority to do so. The existence of a consequence does not prove its justice, neither does someone's attempt to avoid it prove that they believe the consequence is deserved. Someone who thoughtfully uses AI, while believing that they remain responsible for their final work, might genuinely reject the moral framework that turns their use into evidence of intellectual fraud, hence deciding to conceal their use.
This is where watermarking becomes interesting to me. A watermark can make a fact about provenance more detectable, but it cannot tell us what that fact should mean. It cannot establish that the use was intellectually dishonest or that a particular consequence should follow. The technology makes something legible, but the surrounding moral order decides what to do with that legibility. And that is how I somehow arrived at criminals.
A crime describes prohibited conduct; a criminal describes a person through their relationship to that conduct. While the distinction may initially seem pedantic, it is worth noting what happens during that grammatical transformation. It is true that we need categories through which to identify wrongdoing. A society unable to distinguish genuine harm from ordinary conduct would hardly be one I wanted to live in. But, identifying conduct and deciding what kind of person committed it are not exactly the same operation. The category can begin acquiring moral momentum. Someone commits a crime and receives a consequence intended to respond to it. Then another consequence follows. And another. Employment becomes harder to obtain. Housing becomes more difficult. Social exclusion becomes easier to justify. Humiliation feels less troubling. Suspicion follows the person into circumstances increasingly distant from the original act. Each consequence may have its own legitimate justification, but the existence of the category can make us less inclined to demand one.
They're a criminal. Two words can begin paying the moral bill for an astonishing amount of subsequent treatment. This is where I want to interrupt the momentum with questions that appear almost annoyingly elementary: what exactly did this person do? What harm occurred? What is this punishment trying to accomplish? Why did the act become prohibited? What assumptions about the person are entering through the label? Are we responding to conduct or to a category of person? Would we defend this particular treatment if the morally loaded noun were removed from the sentence? The point of these questions is not to make wrongdoing disappear, but rather to prevent one justified consequence from becoming an inexhaustible source of justification for everything that follows.
I have become particularly suspicious of our appetite for cathartic justice. There are acts that produce such anger, disgust, or horror that seeing the person responsible suffer can itself feel like moral restoration. While I understand the feeling, I have also learned enough about my own fallibility to know that the intensity of a feeling does not make the feeling true, much less an appropriate instrument for determining someone else's fate. Our desire to see someone suffer is not evidence of how much suffering a just system should impose. This does not make the desire meaningless, since moral intuition often arrives before argument. Sometimes, something happens and all we possess initially is the stubborn feeling that this is wrong. However, intuition has also sustained terrible prejudice. Human beings have experienced disgust, fear, and moral certainty toward people who did nothing to deserve them. Feeling cannot, therefore, become our final court of appeal simply because the feeling is sincere.
It can, however, earn an investigation: is our reaction proportionate to what occurred? What exactly are we responding to? What harm are we trying to prevent? What assumptions have we smuggled into the category? Why does this person's suffering feel deserved to me? Would I reach the same conclusion if I knew more about them? Am I asking for accountability or catharsis? The rule deserves reasons, the discomfort deserves attention, but neither deserves automatic victory. This is where I find the category of "criminal" even more difficult to treat as a neutral description. Its usefulness is obvious when it helps us identify conduct requiring a social response. Its danger appears when the category begins to flatten and exhaust the person inside it.
A person who commits a terrible act remains a person; they still possesses an interior life. They have motives, fears, memories, relationships, contradictions, reasons, perhaps remorse, perhaps none, and an experience of the world unavailable to everyone standing outside of them. Acknowledging this complex reality does not vindicate them, as explanation is not exoneration, but it prevents guilt from becoming total ontology. Someone can be guilty and more than their guilt. Perhaps this is the minimum burden of treating another person as fully human: acknowledging interiority even when we have every emotional reason to stop imagining it. And when we cannot do that, as I sometimes cannot, perhaps integrity requires acknowledging something about ourselves rather than transforming our limitation into a theory about the other person. "I cannot extend empathy to this person" is different from "this person is undeserving of human consideration"; "I cannot understand how someone could do this" is different from "there is nothing inside them worth understanding". The first admits the limits of my capacity, while the second converts those limits into facts about someone else's humanity. Sometimes the most responsible judgment we can make is a judgment about our own capacity to judge.
Not everyone has what it takes to remain that clear-eyed about highly charged subjects. I don't always. Victims, families, communities, and a frightened public can have completely understandable reasons for wanting outcomes shaped by anger or pain. But, the institutions entrusted with imposing consequences should exist partly because consequential judgment requires distinctions that ordinary human emotion can make difficult to preserve: guilt from personhood, accountability from vengeance, protection from humiliation, and punishment from permanent moral contamination. Authority should require more than moral conviction.
This is one reason history makes me uneasy about the confidence with which categories are wielded. Human societies have repeatedly constructed morally certain explanations for why particular people or behaviors deserved exclusion and diminished consideration, only for later generations to look back and find the certainty itself horrifying. The histories differ too greatly for me to flatten them into one story, and yesterday's injustice does not establish that whatever we condemn today will become tomorrow's vindicated cause. But, it should at least teach us that the confidence of the classifier does not guarantee the justice of the classification.
There have always been people living underneath categories that respectable society considered obvious while knowing, from inside their own lives, that something about the category failed to capture them. Before society possessed or accepted the framework through which their objection could become authoritative, they may have had little more than the knowledge that the treatment felt wrong. This is where I think about the people who are always on the right side of history. I don't mean people who are always morally right; I doubt such people exist. I mean people who have had relatively little experience of respectable society confidently applying the wrong category to them. They have rarely had to occupy the strange position of hearing institutions, intellectual authorities, neighbors, colleagues, or perhaps an entire culture explain what people like them are, while knowing from the inside that something essential has been lost in the description.
Perhaps one of the things privilege can buy is not merely protection from punishment, but confidence in the categories through which punishment is distributed. If the dominant classifications have generally recognized you well enough, you may have less experiential reason to distrust the movement from, "We have classified this behavior" to "Therefore what happens next is deserved". Moral certainty can feel primarily like courage when you have rarely been its mistaken object. Someone who has lived underneath a false classification, on the other hand, may hear something else in the word "deserved". This does not mean that marginalization produces universal wisdom, as someone can know intimately what it feels like to be misclassified in one context and become a ruthless classifier in another. Experiencing prejudice does not inoculate us against prejudice. Being the underdog once does not prevent us from becoming confident in our categories when somebody else occupies the bottom. Perhaps that is another reason categories deserve scrutiny: it is remarkably easy to forget their fallibility once we are holding the label instead of living underneath it.
Which brings me back to AI. Someone may use AI while sincerely believing they remain intellectually responsible for the resulting work. They may have challenged the model and supplied the reasoning behind a piece. They may also know that none of those details will survive the category awaiting them. AI user. Once that fact becomes legible, the conduct can become identity. And concealment then performs a strange moral trick. "Why hide it if you did nothing wrong?" becomes evidence that the person themselves accepts the legitimacy of the category. But, perhaps what they know is not that they did something wrong, but rather that once this particular fact becomes visible, other people may use it to decide what kind of person they are, and set in motion consequences that no longer have any real connection to the original act. This does not establish that concealment is justified, as there are contexts in which AI disclosure may be genuinely owed. But, the existence of those contexts does not resolve every other one, and concealment cannot relieve us of the obligation to establish what was actually owed before declaring the person guilty of refusing to provide it.
Watermarking therefore interests me less because it can reveal AI involvement than because it reduces a person's control over when that involvement becomes legible. The technology reveals a fact into a culture that has already spent years deciding what that fact might mean. And, once a morally contaminated category becomes technically easier to detect, I think we should become more invested, not less, in the consequences attached to detection. Perhaps that is what I wanted from the question "What is a criminal?" all along. Not the abolition of categories, nor the fantasy that nobody does anything wrong, or a world where murder and theft become matters of personal interpretation, but a honest reckoning of what we allow these categories to do to other humans. A person may deserve accountability, but that does not require pretending that the act has swallowed the person whole. And when one is too angry, frightened, disgusted, or morally certain to remember that, perhaps the most responsible thing they can do is not demand that the category become harsher. Perhaps it is to admit that, for the moment, they should not be the one deciding what happens next.