Shared posts

28 Mar 18:37

"Three-person embryos" are now possible

by Minnesotastan
The [British]government is considering whether to propose legal changes that would allow radical new treatments for families at risk of incurable genetic diseases that involve the creation of so-called "three-person embryos".

A national consultation released on Wednesday by the UK's fertility watchdog found public support for techniques that involve introducing DNA from a third person to embryos which could prevent mothers from passing on devastating diseases, such as muscular dystrophy, to their children.

If ministers and MPs give the procedures the green light, Britain would become the first country to offer treatments that lead to children being born with DNA from three people: their parents and a woman donor. The amount of DNA from the donor is tiny compared with the parents...
Scientists have developed two techniques to prevent faulty mitochondria being passed on to children. Known as maternal spindle transfer and pronuclear transfer, they both involve transferring the genetic material from the parents into an egg donated by a healthy woman.
The treatment is controversial on several grounds, not least that the genetic modifications in the embryo pass down to all future generations. The techniques have never been tried in humans, but have worked in animal studies.
Further details at The Guardian.
27 Mar 12:16

Opinions, Morals and What Science Could but Shouldn’t Tell Us

by Sabine Hossenfelder
Headache. Image source: Mupso.
In an opinion piece from December, Brian Cox and Robin Ince argued that opinion must be separated from science when it comes to policy decisions:
“[T]here must be a place where science stops and politics begins, and this border is an extremely complex and uncomfortable one. Science can’t tell us what to do… The choice of policy response itself is not a purely scientific question, however, because it necessarily has moral, geopolitical and economic components.”
I used to say the same, that politics unfortunately mixes up scientific questions with unscientific ones, and that informed decision making requires us to first distinguish these. But then I went down a windy road trying to understand where science ends and where decision making begins. This eventually lead to my paper on the measurement of happiness. It also lead me to the conviction that the “extremely complex and uncomfortable border” doesn’t exist. Cox and Ince come to the right conclusion, but for the wrong reasons.

What is and what isn’t in the realm of science, and what is the role of science is in our political system are questions I care about deeply. And so I could not avoid noticing Sean Carroll and Lubos Motl recently discussed whether morals can be reduced to science. They come down, in rare agreement, on the side of “no”. It’s a variant of the “boundary” Cox and Ince touched on, so let us see what they had to say.

Sean and Lubos start by elaborating on what is and what isn’t a scientific statement. A scientific statement, they say, is one that could be false and whose truth value could at least in principle be empirically evaluated. The problem is then that the statement that morals can’t be reduced to science itself isn’t scientific. It isn’t because a definition for “moral” is lacking. Then, all answers to this question are just opinion so why bother with it? Lubos alludes to this by saying that whenever one could answer the question one way or the other somebody might just change the definition of moral
“Imagine that you find some quantity M encoded in the equations of M(orality)-theory in the future and you will claim that it measures morality… The problem is that even with this nice and well-defined formula, one may always legitimately refuse such a measure of morality and choose a completely different one.”
This lack of proper definition is an example for what I complained about in my recent post, that many philosophical questions are a waste of time if one doesn’t know what one is talking about to begin with. So let’s not debate the meaning of words and instead identify the real issue behind it.

What people really want to know is where science leaves them the freedom to make decisions. That’s why they are looking for a border between scientific and unscientific questions, the former can be answered by science, the latter presumably can only be answered by humans. In other words, they’re asking for their space to exercise free will.

Free will is an illusion that people hold on to quite stubbornly and that they protect vehemently, so the debate about the unscientificness of morals shouldn’t come as a surprise. The thought that science might tell people what they should or shouldn’t do is a great threat to free will, one that gets addressed in a forward defense. But that’s a misunderstanding. Science has never and will never tell anybody what should or shouldn’t be done because “should” is another one of these ill-defined words. “Should” implicitly necessitates a goal or a purpose.

“Science can’t tell us what to do”, as Cox and Ince write - correctly. But science can in principle tell us what we do. To understand how let’s have a look at what people mean when they refer to “morals” or “values”.

Humans are self-aware complex systems that have to process a lot of information to make informed decisions. Human self-awareness however is limited. We are not normally aware how the detailed processes of our thoughts proceed. In fact recent research in neuroscience seems to show that what we think of as “I” is primarily an aggregating mechanism of various deeper level systems whose detailed procedures the “I” does not normally take note of.

Thus, “we” don’t consciously know the details of how we make decisions. Moreover, a central element of human decision making is ignorance and oversimplification. The one thing that the human brain is really good at is energy efficiency. Which is why the default is to avoid thinking if unnecessary.

What we do instead of monitoring all that information from the input that we receive is learning to construct models of behavior that make use of simplified patterns and categories. Then we explain our decisions and those of others in terms of these simplified patterns. You chose this job because independence is important to you. You think polygamy is immoral and should be punished. These are rough summaries of longwinded thought processes which made use of experience, evolutionary traits, and random noise. They classify decisions in values like “independence” or morals like “faithfulness.”

Morals and values are thus just categories that people use to classify and explain the way they make decisions. Over time, using these simplified models, the higher level “I” system becomes good at predicting what will happen, and interprets this as an exercise of “free will.”

That having been said, if you believe in reductionism, morals and values are just emergent patterns in highly complex systems. It is clearly impractical and anyway presently impossible, but in principle one could define morals in this way. Imagine you’d do this. Now you have a definition for moral. An individual one, one that depends on cultural history as well as genetic ancestry. Here you have it. These are your morals.

You might then go and say that’s not what you mean with moral. And that would be fine with me because I don’t want to argue about words, so just call these emerging patterns something else. The point is that they’re what people make use of when they make decisions, and recall that this is the question we really want to address: What decisions are humans free to make because they’re allegedly unscientific?

If you have such a definition for morals then would science then tell you what you should do? No. It would in the best case simply tell you what you do. The best case being one in which scientists would be able to construct a complete model for human behavior. Depending on your attitude you might call that the worst case.

But while in principle possible, it is questionable that such a model is feasible to construct at all. It seems plausible to me that the process of thought is irreducible in the sense that if you tried to predict it you’d have to create an almost perfect copy of the original system and watch it in real time, in which case you’d just duplicate rather than predict decisions.

In other words, while the “border” between scientific and unscientific questions does not exist in principle, it does exist in practice. And it’s located where our ability to model complex systems ends, an end that might shift somewhat in the future but quite possibly will never entirely recede. The best way we presently know to find out what decisions humans make is to ask them. The best way we know to find out what the global climate does is not to ask humans but a computer model.

What does this have to do with happiness? Well, striving to achieve happiness is a human universal, so much so that you might want to raise the maximization of well-being of conscious beings to a universal goal. Having defined such a goal it would fill in the blank of the “purpose” and the “should” that was previously missing, or at least it seems so.

The problem is however that happiness is a byeffect of natural selection, it’s a simplified response to behavior that has in the past been beneficial for reproduction. Elevating happiness to an end unto itself is a circular definition of purpose, it’s fundamentally meaningless. Which is why, in my paper I argued we should forget about trying to define happiness and its maximization as a proxy to understand human behavior. Instead we should look for a properly defined quantity that has predictive power to describe the evolution of our economic, politic, and social systems, and the suggestion I made was maximizing the number of possible decisions that we (think we can) make. Which might or might not be correct. A scientific question that’s waiting to be answered.

Summary: Ill-defined questions are unscientific, but uninterestingly so. Once a question is well-defined science is in principle able to answer it, but not necessarily in practice. A scientific definition for morals might exist, but quite plausibly we will never be able to construct it. And even if we could, it wouldn’t tell us what should be done, but simply what is done. Opinion begins where our ability to model complex systems ends. This border will inevitably shift over time and it’s this “shift” that makes it uncomfortable. And no, I don’t believe in free will.
06 Mar 22:46

Police hunt Boston hater who punches out man, who then gets run over by a cab

by adamg

UPDATE: A law-enforcement source and a fan say the victim in the attack was One Armed Push-Up Guy, well known outside Boylston Street bars and in the Fenway for doing one-armed pushups for a fee, like this:

Boston Police report that around 2 a.m., a guy standing in the middle of Boylston Street outside Whiskey's, 885 Boylston St., yelled "Fuck Boston and fuck your businesses!" and then turned around and punched a nearby man, who fell unconscious into the street:

Witnesses reported a cab traveling inbound on Boylston Street ran over the victim while he was unconscious in the road and dragged him under the vehicle.

The screaming puncher then ran inbound on Boylston and disappeared down Fairfield Street, police say. He's described as 30 to 32, 5'10" and clad in a dark jacket with a fur-lined hood.

The cab driver told police he did not realize he'd hit somebody and stopped only when bystanders started motioning to him. The victim was taken to Mass. General and is expected to survive, police say.

If you know the guy, contact D-4 detectives at 617-343-5619 or the anonymous tip line by calling 800-494-TIPS or texting TIP to CRIME (27463).

06 Mar 22:18

Mismatch between the Na+ flux and spike initiation [Neuroscience]

by Baranauskas, G., David, Y., Fleidervish, I. A.
Nosimpler

what the hell.

It is widely believed that, in cortical pyramidal cells, action potentials (APs) initiate in the distal portion of axon initial segment (AIS) because that is where Na+ channel density is highest. To investigate the relationship between the density of Na+ channels and the spatiotemporal pattern of AP initiation, we simultaneously...
06 Mar 18:09

Blowgun darts - a new public menace?

by Minnesotastan

The image comes from a Reddit post, with this explanation:
Was walking on the local trail today, going through the forest part. When suddenly I feel a very sharp pain in my arm, like a shot. Look down and there is a long dart, from maybe a blow gun of sorts? I don't know but scared the shit out of me, I looked and could see no one.
A simple Google Image search yields a wide array of photos of this weaponry.  The Reddit thread focuses on potential health consequences (probably minimal other than at the site of entry) and the range (up to about 100 feet, if you want to locate and try to wreak retribution on the malefactor).
06 Mar 15:27

Spivak on Category Theory

by willerton
Nosimpler

I met with David last year and it was one of the more enlightening experiences I've had. Maybe we should do a study group on this?

MathML-enabled post (click for more details).

Guest post by Bruce Bartlett

We know about Category Theory for Mathematicians, we’ve all read Category Theory for Physicists, and we also know about Category Theory for Computer Scientists, and we’ve even seen the videos.

But how about Category Theory for Scientists? I spotted this on the arXiv listings.

David Spivak, Category Theory for Scientists.

Abstract: There are many books designed to introduce category theory to either a mathematical audience or a computer science audience. In this book, our audience is the broader scientific community. We attempt to show that category theory can be applied throughout the sciences as a framework for modeling phenomena and communicating results. In order to target the scientific audience, this book is example-based rather than proof-based. For example, monoids are framed in terms of agents acting on objects, sheaves are introduced with primary examples coming from geography, and colored operads are discussed in terms of their ability to model self-similarity.

pic of coverMathML-enabled post (click for more details).

I’m afraid this little post is just a shout-out as I’ve only hurriedly browsed through the pages.

Towards the end of the book he gets to sheaves; he is certainly an expert on these as his PhD thesis was on derived smooth manifolds). His motivating example is stitching together pictures of the night sky, which I thought was really cool:

pic of stars

Paging through, I see the Yoneda lemma only gets a small paragraph, with a reference to Mac Lane. I’m kind of sad about that, since I do regard it as the fundamental theorem of category theory. Too bad.

11 Feb 18:16

Torture in a Just World

by Alex Tabarrok

If the world is just, only the guilty are tortured. So believers in a just world are more likely to think that the people who are tortured are guilty. Perhaps especially so if they experience the torture closely and so feel a greater need to overcome cognitive dissonance. On the other hand, those farther away from the experience of torture may feel less need to justify it and they may be more likely to identify the tortured as victims. The theory of moral typecasting suggests that victims are also more likely to be seen as innocents (a la Jesus).

The theory is tested in a lab setting by Gray and Wegner. Experimental subjects are told that “Carol”, really a confederate, may have lied about a dice roll and that stress often encourages people to admit guilt. Subjects then listen to a torture session as Carol’s hand is plunged into a bucket of ice water for 80s. Subjects are then asked how likely is it that the torture victim was lying (1 to 5 with 5 being extremely likely). There are two intervention variables: 1) some of the subjects meet the torture victim before she is tortured, this is the close condition and some do not (distance condition) and 2) in some torture sessions the victim evinces pain (pain) and in others not (no pain). The key figure is shown below:

torturegraphThe most striking result is that in the close condition, the evincing of pain was associated with an increased judgment of guilt, consistent with torture causing cognitive dissonance which is relieved by a judgment of guilt (restoring the just world). But in the distance condition, the evincing of pain was associated with a decreased judgement of guilt, consistent with pain increasing the identification of the tortured as a victim and therefore innocent (a la moral typecasting).

Closeness in the experiment was reasonably literal but may also be interpreted in terms of identification with the torturer. If the church is doing the torturing then the especially religious may be more likely to think the tortured are guilty. If the state is doing the torturing then the especially patriotic (close to their country) may be more likely to think that the tortured/killed/jailed/abused are guilty. That part is fairly obvious but note the second less obvious implication–the worse the victim is treated the more the religious/patriotic will believe the victim is guilty.

The theory has interesting lessons for entrepreneurs of social change. Suppose you want to change a policy such as prisoner abuse (e.g. Abu Ghraib) or no-knock police raids or the war on drugs or even tax policy. Convincing people that the abuse is grave may increase their belief that the victim is guilty. Instead, you want to do one of two things. Among the patriotic you may want to sell the problem as a minor problem that We Can Fix – making them feel good about both the we and the fixing. Or, you may want to create distance – The problem is bad and THEY are the cause. People in the North, for example, became more concerned about slavery once the US became us and them.

I think research in moral reasoning is important because understanding why good people do evil things is more important than understanding why evil people do evil things.

11 Feb 18:14

Occupy HSBC: Valentine’s Day protest at noon #OWS

by Cathy O'Neil, mathbabe

Protest with #OWS Alternative Banking Group

I’m writing to invite you to a protest against mega-bank HSBC at noon on Valentine’s Day (Thursday) starting on the steps of the New York Public Library at 42nd and 5th. Details are here but it’s the big green box on the map on the Fifth Avenue side:

Screen Shot 2013-02-11 at 7.13.43 AM

Why are we protesting?

Like you, I’m sure, I’d like nothing more than to stop worrying about shit that goes on in our country’s banks.

We have better things to do with out time than to get annoyed over enormous bonuses being given to idiots for their repeated failures. We’re frankly exhausted from the outrage.

I mean, the average person doesn’t have a job where they get an $11 million bonus instead of a $22 million dollar bonus when they royally screw up. Outside the surreal realm of international banking, the normal response to screw-ups on that level is to get fired.

You might expect a company that has been caught criminally screwing minorities out of fair contracts might be at risk of being closed down, but in this day and age you’d know that big banks, or TIBACO (too interconnected, big, and complex to oversee) institutions, as we in Alt Banking like to call them, are immune to such action.

There’s a clear evolving standard of treatment in the banking sector when it comes to criminal activity:

  • the powers that be (SEC, DOJ, etc.) make a huge production over the severity of the fine,
  • which is large in dollar amounts but
  • usually represents about 10% of the overall profit the given banks made during their exploit.
  • Nobody ever goes to jail, and
  • the shareholders pay the fine, not the perpetrators.
  • The perps get somewhat diminished bonuses. At worst.

The bottomline: we have an entire class of citizens that are immune to the laws because they are considered too important to our financial stability.

scarface_customer_criminal

But why HSBC?

HSBC is a perfect example of this. An outrageous example.

HSBC didn’t get a bailout in 2008 like many other banks, even though they were ranked #2 in subprime mortgage lending. But that’s not because they didn’t lose money – in fact they lost $6 billion but somehow kept afloat.

And now we know why.

Namely, they were money-laundering, earning asstons by  facilitating drugs and terrorism. This was blood money, make no mistake, and it went directly into the pockets of HSBC bankers in the form of bonuses.

When this years-long criminal mafia activity was discovered, nothing much happened beyond a fine, as per usual. Well, to be honest, they were fined $1.9 billion dollars, which is a lot of money, but is only 5 weeks of earnings for the mammoth institution – depending on the way you look at it, HSBC is the 2nd largest bank in the world.

dirty_money_HSBC

Too big to jail

And that’s when “Too big to fail” became “Too big to jail.” Even the New York Times was outraged. From their editorial page:

Federal and state authorities have chosen not to indict HSBC, the London-based bank, on charges of vast and prolonged money laundering, for fear that criminal prosecution would topple the bank and, in the process, endanger the financial system. They also have not charged any top HSBC banker in the case, though it boggles the mind that a bank could launder money as HSBC did without anyone in a position of authority making culpable decisions.

Clearly, the government has bought into the notion that too big to fail is too big to jail. When prosecutors choose not to prosecute to the full extent of the law in a case as egregious as this, the law itself is diminished. The deterrence that comes from the threat of criminal prosecution is weakened, if not lost.

HSBC_AD

National Threat

You may recall that there was an extensive FBI investigation of OWS before Zuccotti Park was even occupied.

Ironic? As the Village Voice said, “apparently non-violent demonstration against corrupt banking is subject to more criminal scrutiny than actual corrupt banking.”

Question for you: which is the bigger national security threat, OWS or HSBC?

drug_lord_bank_of_choice_HSBC

We demand

HSBC needs its license revoked, and there need to be prosecutions. Those who are guilty need to be punished or else we have an official invitation to criminal acts by bankers. We simply can’t live in a country which rewards this kind of behavior.

Mind you, this isn’t just about HSBC. This is about all the megabanks. Citi or BoA are exempt from prosecution, too. Our message needs to be “break up the megabanks”.

I’ll end with what Matt Taibbi had to say about the HSBC settlement:

On the other hand, if you are an important person, and you work for a big international bank, you won’t be prosecuted even if you launder nine billion dollars. Even if you actively collude with the people at the very top of the international narcotics trade, your punishment will be far smaller than that of the person at the very bottom of the world drug pyramid. You will be treated with more deference and sympathy than a junkie passing out on a subway car in Manhattan (using two seats of a subway car is a common prosecutable offense in this city). An international drug trafficker is a criminal and usually a murderer; the drug addict walking the street is one of his victims. But thanks to Breuer, we’re now in the business, officially, of jailing the victims and enabling the criminals.

Join us on Valentine’s Day at noon on the steps of the New York Public Library and help us Occupy HSBC. Please redistribute widely!

cherub


05 Feb 21:20

sangfroid

Merriam-Webster's Word of the Day for February 05, 2013 is:

sangfroid • \SAHNG-FRWAH\  • noun
: self-possession or imperturbability especially under strain

Examples:
The lecturer's sangfroid never faltered, even in the face of some tough questions from the audience.

"Daniel Craig portrays a vulnerability far removed from the glib sangfroid of his celluloid predecessors and has retired to an exotic bolt-hole after he is assumed to have died during a botched operation." — From a movie review by Des O' Neill in the Irish Times, January 2, 2013

Did you know?
If you're a lizard, "cold-blooded" means your body temperature is strongly influenced by your environment. If you're an English-speaking human, it means you are callous and unfeeling. If you're a French speaker, it means that you're calm, cool, and collected in stressful situations. By the mid-1700s, English speakers had already been using "cold-blooded" for more than a century, but they must have liked the more positive spin the French put on having "cold blood" because they borrowed the French "sang-froid" (literally, "cold blood") for someone who is imperturbable under strain. The French term, by the way, developed from the Latin words "sanguis" ("blood") and "frigidus" ("cold").

24 Jan 21:07

"Fire in the Blood": Millions Die in Africa After Big Pharma Blocks Imports of Generic AIDS Drugs

by mail@democracynow.org (Democracy Now!)
Fireintheblood

The new documentary, "Fire in the Blood," examines how millions have died from AIDS because big pharmaceutical companies and the United States have refused to allow developing nations to import life-saving generic drugs. The problem continues today as the World Trade Organization continues to block the importation of generic drugs in many countries because of a trade deal known as the TRIPS Agreement. We’re joined by the film’s director, Dylan Mohan Gray, and Ugandan AIDS doctor Peter Mugyenyi, who was arrested for trying to import generic drugs, and is recognized as one of the world’s foremost specialists and researchers in the field of HIV/AIDS. [includes rush transcript]

12 Jan 23:44

Mike Izbicki: The categorical distribution’s algebraic structure

histogram of simonDist The categorical distribution is the main distribution for handling discrete data. I like to think of it as a histogram.  For example, let’s say Simon has a bag full of marbles.  There are four “categories” of marbles—red, green, blue, and white.  Now, if Simon reaches into the bag and randomly selects a marble, what’s the probability it will be green?  We would use the categorical distribution to find out.

In this article, we’ll go over the math behind the categorical distribution, the algebraic structure of the distribution, and how to manipulate it within Haskell’s HLearn library.  We’ll also see some examples of how this focus on algebra makes HLearn’s interface more powerful than other common statistical packages.  Everything that we’re going to see is in a certain sense very “obvious” to a statistician, but this algebraic framework also makes it convenient.  And since programmers are inherently lazy, this is a Very Good Thing.

Before delving into the “cool stuff,” we have to look at some of the mechanics of the HLearn library.

Preliminaries

The HLearn-distributions package contains all the functions we need to manipulate categorical distributions. Let’s install it:

$ cabal install HLearn-distributions

We import our libraries:

>import Control.DeepSeq
>import HLearn.Algebra
>import HLearn.Gnuplot.Distributions
>import HLearn.Models.Distributions

We create a data type for Simon’s marbles:

>data Marble = Red | Green | Blue | White
>    deriving (Read,Show,Eq,Ord)

marbles

The easiest way to represent Simon’s bag of marbles is with a list:

>simonBag :: [Marble]
>simonBag = [Red, Red, Red, Green, Blue, Green, Red, Blue, Green, Green, Red, Red, Blue, Red, Red, Red, White]

And now we’re ready to train a categorical distribution of the marbles in Simon’s bag:

>simonDist = train simonBag :: Categorical Marble Double

We can load up ghci and plot the distribution with the conveniently named function plotDistribution:

ghci> plotDistribution (plotFile "simonDist") simonDist

This gives us a histogram of probabilities:

marbles trained into categorical

In the HLearn library, every statistical model is generated from data using either train or train’.  Because these functions are overloaded, we must specify the type of simonDist so that the compiler knows which model to generate. Categorical takes two parameters. The first is the type of the discrete data (Marble). The second is the type of the probability (Double). We could easily create Categorical distributions with different types depending on the requirements for our application. For example:

>stringDist = train (map show simonBag) :: Categorical String Float

This is the first “cool thing” about Categorical:  We can make distributions over any user-defined type.  This makes programming with probabilities easier, more intuitive, and more convenient.  Most other statistical libraries would require you to assign numbers corresponding to each color of marble, and then create a distribution over those numbers.

Now that we have a distribution, we can find some probabilities. If Simon pulls a marble from the bag, what’s the probability that it would Red?

We can use the pdf function to do this calculation for us:

ghci> pdf simonDist Red
0.5626
ghci> pdf simonDist Blue
0.1876
ghci> pdf simonDist Green
0.1876
ghci> pdf simonDist White
6.26e-2

If we sum all the probabilities, as expected we would get 1:

ghci> sum $ map (pdf simonDist) [Red,Green,Blue,White]
1.0

Due to rounding errors, you may not always get 1. If you absolutely, positively, have to avoid rounding errors, you should use Rational probabilities:

>simonDistRational = train simonBag :: Categorical Marble Rational

Rationals are slower, but won’t be subject to floating point errors.

This is just about all the functionality you would get in a “normal” stats package like R or NumPy. But using Haskell’s nice support for algebra, we can get some extra cool features.

Semigroup

First, let’s talk about semigroups. A semigroup is any data structure that has a binary operation (<>) that joins two of those data structures together. The categorical distribution is a semigroup.

Don wants to play marbles with Simon, and he has his own bag. Don’s bag contains only red and blue marbles:

>donBag = [Red,Blue,Red,Blue,Red,Blue,Blue,Red,Blue,Blue]

We can train a categorical distribution on Don’s bag in the same way we did earlier:

>donDist = train donBag :: Categorical Marble Double

In order to play marbles together, Don and Simon will have to add their bags together.

>bothBag = simonBag ++ donBag

Now, we have two options for training our distribution. First is the naive way, we can train the distribution directly on the combined bag:

>bothDist = train bothBag :: Categorical Marble Double

This is the way we would have to approach this problem in most statistical libraries. But with HLearn, we have a more efficient alternative. We can combine the trained distributions using the semigroup operation:

>bothDist' = simonDist <> donDist

Under the hood, the categorical distribution stores the number of times each possibility occurred in the training data.  The <> operator just adds the corresponding counts from each distribution together:

semigroup and bothDist

This method is more efficient because it avoids repeating work we’ve already done. Categorical’s semigroup operation runs in time O(1), so no matter how big the bags are, we can calculate the distribution very quickly. The naive method, in contrast, requires time O(n). If our bags had millions or billions of marbles inside them, this would be a considerable savings!

We get another cool performance trick “for free” based on the fact that Categorical is a semigroup: The function train can be automatically parallelized using the higher order function parallel. I won’t go into the details about how this works, but here’s how you do it in practice.

First, we must show the compiler how to resolve the Marble data type down to “normal form.” This basically means we must show the compiler how to fully compute the data type. (We only have to do this because Marble is a type we created.  If we were using a built in type, like a String, we could skip this step.) This is fairly easy for a type as simple as Marble:

>instance NFData Marble where
>    rnf Red   = ()
>    rnf Blue  = ()
>    rnf Green = ()
>    rnf White = ()

Then, we can perform the parallel computation by:

>simonDist_par = parallel train simonBag :: Categorical Marble Double

Other languages require a programmer to manually create parallel versions of their functions. But in Haskell with the HLearn library, we get these parallel versions for free! All we have to do is ask for it!

Monoid

A monoid is a semigroup with an empty element, which is called mempty in Haskell. It obeys the law that:

M <> mempty == mempty <> M == M

And it is easy to show that Categorical is also a monoid. We get this empty element by training on an empty data set:

mempty = train ([] :: [Marble]) :: Categorical Marble Double

The HomTrainer type class requires that all its instances also be instances of Monoid. This lets the compiler automatically derive “online trainers” for us. An online trainer can add new data points to our statistical model without retraining it from scratch.

For example, we could use the function add1dp (stands for: add one data point) to add another white marble into Simon’s bag:

>simonDistWhite = add1dp simonDist White

This also gives us another approach for our earlier problem of combining Simon and Don’s bags. We could use the function addBatch:

>bothDist'' = addBatch simonDist donBag

Because Categorical is a monoid, we maintain the property that:

bothDist == bothDist' == bothDist''

Again, statisticians have always known that you could add new points into a categorical distribution without training from scratch.  The cool thing here is that the compiler is deriving all of these functions for us, and it’s giving us a consistent interface for use with different data structures.  All we had to do to get these benefits was tell the compiler that Categorical is a monoid.  This makes designing and programming libraries much easier, quicker, and less error prone.

Group

A group is a monoid with the additional property that all elements have an inverse. This lets us perform subtraction on groups.  And Categorical is a group.

Ed wants to play marbles too, but he doesn’t have any of his own. So Simon offers to give Ed some of from his own bag. He gives Ed one of each color:

>edBag = [Red,Green,Blue,White]

Now, if Simon draws a marble from his bag, what’s the probability it will be blue?

To answer this question without algebra, we’d have to go back to the original data set, remove the marbles Simon gave Ed, then retrain the distribution. This is awkward and computationally expensive. But if we take advantage of Categorical’s group structure, we can just subtract directly from the distribution itself. This makes more sense intuitively and is easier computationally.

>simonDist2 = subBatch simonDist edBag

This is a shorthand notation for using the group operations directly:

>edDist = train edBag :: Categorical Marble Double
>simonDist2' = simonDist <> (inverse edDist)

The way the inverse operation works is it multiplies the counts for each category by -1. In picture form, this flips the distribution upside down:

edDist inversification

Then, adding an upside down distribution to a normal one is just subtracting the histogram columns and renormalizing:

Simon substraction edDist

Notice that the green bar in edDist looks really big—much bigger than the green bar in simonDist.  But when we subtract it away from simonDist, we still have some green marbles left over in simonDist2.  This is because the histogram is only showing the probability of a green marble, and not the actual number of marbles.

Finally, there’s one more crazy trick we can perform with the Categorical group.  It’s perfectly okay to have both positive and negative marbles in the same distribution.  For example:

ghci> plotDistribution (plotFile "mixedDist") (edDist <> (inverse donDist))

results in:

mixedDist-300

Most statisticians would probably say that these upside down Categoricals are not “real distributions.” But at the very least, they are a convenient mathematical trick that makes working with distributions much more pleasant.

Module

Finally, an R-Module is a group with two additional properties. First, it is abelian. That means <> is commutative. So, for all a, b:

a <> b == b <> a

Second, the data type supports multiplication by any element in the ring R. In Haskell, you can think of a ring as any member of the Num type class.

How is this useful?  It let’s “retrain” our distribution on the data points it has already seen.  Back to the example…

Well, Ed—being the clever guy that he is—recently developed a marble copying machine. That’s right! You just stick some marbles in on one end, and on the other end out pop 10 exact duplicates. Ed’s not just clever, but pretty nice too. He duplicates his new marbles and gives all of them back to Simon. What’s Simon’s new distribution look like?

Again, the naive way to answer this question would be to retrain from scratch:

>duplicateBag = simonBag ++ (concat $ replicate 10 edBag)
>duplicateDist = train duplicateBag :: Categorical Marble Double

Slightly better is to take advantage of the Semigroup property, and just apply that over and over again:

>duplicateDist' = simonDist2 <> (foldl1 (<>) $ replicate 10 edDist)

But even better is to take advantage of the fact that Categorical is a module and the (.*) operator:

>duplicateDist'' = simonDist2 <> 10 .* edDist

In picture form:

module example

Also notice that without the scalar multiplication, we would get back our original distribution:

module example-mod

Another way to think about the module’s scalar multiplication is that it allows us to weight our distributions.

Ed just realized that he still needs a marble, and has decided to take one.  Someone has left their Marble bag sitting nearby, but he’s not sure whose it is.  He thinks that Simon is more forgetful than Don is, so he assigns a 60% probability that the bag is Simon’s and a 40% probability that it is Don’s.  When he takes a marble, what’s the probability that it is red?

We create a weighted distribution using module multiplication:

>weightedDist = 0.6 .* simonDist <> 0.4 .* donDist

Then in ghci:

ghci> pdf weightedDist Red
0.4929577464788732

We can also train directly on weighted data:

>weightedDataDist = train [(0.4,Red),(0.5,Green),(0.2,Green),(3.7,White)] :: Categorical Marble Double

which gives us:

weightedDataDist-300

The Takeaway and next posts

Talking about the categorical distribution in algebraic terms let’s us do some cool new stuff with our distributions that we can’t easily do in other libraries.  None of this is statistically ground breaking. The cool thing is that algebra just makes everything so convenient to work with.

I think I’ll do another post on some cool tricks with the kernel density estimator that are not possible at all in other libraries, then do a post about the category (formal category-theoretic sense) of statistical training methods.  At that point, we’ll be ready to jump into machine learning tasks.  Depending on my mood we might take a pit stop to discuss the computational aspects of free groups and modules and how these relate to machine learning applications.

Sign up for the RSS feed to stay tuned!

12 Jan 23:43

Roman Cheplyaka: Subtractable values are torsors

A group is an abstraction for things that can be «added» and «subtracted» (in some sense).

Most of the groups that programmers deal with are numeric, although some interesting groups arise in graphics and cryptography.

If we only require that things can be added, but not subtracted, we get semigroups and monoids (the difference is that a monoid also has a «zero» element). These are plentiful, with applications ranging from diagrams to regular expressions.

The reason is, of course, that addition is very natural — we can consider any way to combine objects together an addition as soon as it is associative — while subtraction, the inverse of addition, often doesn’t make sense.

Dates

But there are also things that can be subtracted, but not added. The simplest example is dates. You can subtract two dates to get the time interval between them, but you cannot add two dates. It doesn’t make sense to ask what today plus tomorrow is. You could try to add the numeric values corresponding to the dates and convert the answer back to a date, but the result would depend on whether you count from the start of the Common Era, the Unix epoch, or something else.

While adding two dates is not possible, it is possible to add a time interval to a date («five days from today»). This suggests that we should not confound dates and time intervals — they are different types of values.

For example, in the Haskell time package, there are types UTCTime and NominalDiffTime, and the following functions:

addUTCTime :: NominalDiffTime -> UTCTime -> UTCTime
diffUTCTime :: UTCTime -> UTCTime -> NominalDiffTime

Time intervals can also be added and subtracted, as witnessed by the Num NominalDiffTime instance.

Ask mathematics

What is the mathematical notion that characterizes dates? It should involve two sets: the set V of values and the set D of value differences.

The set D must form an additive group, so that we can add and subtract the differences. (The word «additive» means that we assume that the group operation, identity element and inverse operation are called +, 0 and respectively. In group theory it is more common to call them *, 1 and ⁻¹.)

Thus we have the “zero” difference 0∈D and operations +:D×D→D, −:D→D which satisfy the usual group laws:

+/0 (identity element)

0+x=x+0=x

+/+ (associativity)

(x+y)+z=x+(y+z)

+/− (inverse element)

x+(−x)=(−x)+x=0

(We can write x−y as a shorthand for x+(−y).)

We do not require + to be commutative to preserve generality, even though it is commutative for our current example.

We need to be able to add a difference to a value. Thus, we need a function add:D×V→V. We would expect it to play well with the group structure of D:

add/0

add(0,v)=v

add/+

add(x+y,v)=add(x,add(y,v))

An alternative view on these laws is that the curried version of add, curry(add):D→(V→V), is a group homomorphism from D to the group of bijections on V.

The mathematical abstraction corresponding to our add function is an action of group D on the set V. But it doesn’t yet capture the fact that we can subtract two values to get a difference. For that we need a function diff:V×V→D inverse to add:

diff/add

diff(add(x,v),v)=x

add/diff

add(diff(u,v),v)=u

If such a function exists, the set V is called a torsor over the group D.

Thus, it is the mathematical notion of torsor that characterizes subtractable values.

Conal Elliott defines torsors in Haskell under the name of affine spaces.

Other examples

Points

Another simple example of a torsor is points. In linear algebra we get used to conflating points and vectors, because we consider vectors originating at zero.

But the geometrical intuition suggests that vectors do not have to be tied to zero. A vector is a directed segment which can be freely translated (shifted) on the plane (or in the space).

We can subtract two points to get a vector from one point to another. We can add a vector to a point, by translating it so that its origin becomes the given point. The end point of the vector then becomes the result of addition.

Vectors form an additive group, and points form a torsor over the vectors.

Like with dates, results of torsor operations on points and vectors do not depend on (and do not require the existence of) a coordinate system, while the «sum» of two points will change if you change the origin.

Files

There are also some examples that look very much like torsors, but do not satisfy all the laws.

One of such examples is files. The difference of two files is what we call a patch, or a diff. We can diff two files and apply the resulting patch to a third file («cherry-picking»).

Although patches do not form a group, they can be modelled as inverse semigroups. This motivates us to consider torsors based on inverse semigroups rather than on groups.

Haskell dialects

Haskell dialects can be considered as a torsor-like structure over Haskell extensions. We can add extensions to languages, e.g. add(RankNTypes+GADTs−MonomorphismRestriction, Haskell2010).

We can also subtract two languages to find out the difference between them in terms of extensions: diff(Haskell2010, Haskell98)=DoAndIfThenElse+PatternGuards+ForeignFunctionInterface+EmptyDataDecls−NPlusKPatterns.

Extensions do not form a group. The easiest way to see that is to observe that extensions are idempotent, i.e. X+X=X for any extension X. However, in a group there is only one idempotent element — zero.

Extensions do form an inverse semigroup, though. Thus we again arrive at inverse semigroup torsors.

Acknowledgements

A recent discussion with Niklas Broberg about Haskell extensions motivated me to explore these concepts. His HIW’12 lightning talk on Haskell Modular Mindset is also relevant.

Oleksandr Manzyuk kindly pointed me to the concept of a torsor.

31 Dec 06:30

Dense percolation in large-scale mean-field random networks is provably "explosive".

PLoS One. 2012; 7(12): e51883
Veremyev A, Boginski V, Krokhmal PA, Jeffcoat DE

Recent reports suggest that evolving large-scale networks exhibit "explosive percolation": a large fraction of nodes suddenly becomes connected when sufficiently many links have formed in a network. This phase transition has been shown to be continuous (second-order) for most random network formation processes, including classical mean-field random networks and their modifications. We study a related yet different phenomenon referred to as dense percolation, which occurs when a network is already connected, but a large group of nodes must be dense enough, i.e., have at least a certain minimum required percentage of possible links, to form a "highly connected" cluster. Such clusters have been considered in various contexts, including the recently introduced network modularity principle in biological networks. We prove that, contrary to the traditionally defined percolation transition, dense percolation transition is discontinuous (first-order) under the classical mean-field network formation process (with no modifications); therefore, there is not only quantitative, but also qualitative difference between regular and dense percolation transitions. Moreover, the size of the largest dense (highly connected) cluster in a mean-field random network is explicitly characterized by rigorously proven tight asymptotic bounds, which turn out to naturally extend the previously derived formula for the size of the largest clique (a cluster with all possible links) in such a network. We also briefly discuss possible implications of the obtained mathematical results on studying first-order phase transitions in real-world linked systems.

03 Dec 08:21

Eye Movements to Natural Images as a Function of Sex and Personality

by Felix Joseph Mercer Moss et al.
Nosimpler

Apparently men focus on eyes and women focus on chins. Weird.

by Felix Joseph Mercer Moss, Roland Baddeley, Nishan Canagarajah

Women and men are different. As humans are highly visual animals, these differences should be reflected in the pattern of eye movements they make when interacting with the world. We examined fixation distributions of 52 women and men while viewing 80 natural images and found systematic differences in their spatial and temporal characteristics. The most striking of these was that women looked away and usually below many objects of interest, particularly when rating images in terms of their potency. We also found reliable differences correlated with the images' semantic content, the observers' personality, and how the images were semantically evaluated. Information theoretic techniques showed that many of these differences increased with viewing time. These effects were not small: the fixations to a single action or romance film image allow the classification of the sex of an observer with 64% accuracy. While men and women may live in the same environment, what they see in this environment is reliably different. Our findings have important implications for both past and future eye movement research while confirming the significant role individual differences play in visual attention.
02 Dec 20:21

Cryoconite in Birthday Canyon

by Minnesotastan
Nosimpler

Turns out melting glaciers are, in fact, pretty cool to look at. Thanks global warming!

Adam LeWinter on the rim of Birthday Canyon on the Greenland Ice Sheet. The black deposit in the bottom of channel is cryoconite. Birthday Canyon is approximately 150 feet deep.
The photo (credit James Balog/Extreme Ice Survey) comes from a small gallery at The Guardian depicting scenes of glacial melting in the Arctic.  Second gallery.

Blogged for the stark beauty, but it also prompted me to look up what "cryoconite" is.
Cryoconite is powdery windblown dust which is deposited and builds up on snow, glaciers, or icecaps. It contains small amounts of soot which absorbs solar radiation melting the snow or ice beneath the deposit sometimes creating a cryoconite hole. 
Cryoconite holes have been suggested to play important roles in the glacier ecosystems because many kinds of living organisms have been reported from this structure on the glaciers, for example, algae, rotifer, tardigrada, insects and ice worm.
Details about cryoconite granules
13 Nov 22:44

What the Petraeus Investigation Tells Us About Online Surveillance

by J.D. Tuccille

EmailWith regards to the David Petraeus scandal, as you dig through the very human details of a powerful man's dalliance with an attractive woman, an important question should occur to anybody with more than a National Enquirer-level interest in the matter: Wait ... The FBI did all of this digging over some bed-hopping? Yes. Yes, it did. And over at The Guardian, Glenn Greenwald wants to know why more people aren't concerned.

Writes Greenwald:

As is now widely reported, the FBI investigation began when Jill Kelley - a Tampa socialite friendly with Petraeus (and apparently very friendly with Gen. John Allen, the four-star U.S. commander of the war in Afghanistan) - received a half-dozen or so anonymous emails that she found vaguely threatening. She then informed a friend of hers who was an FBI agent, and a major FBI investigation was then launched that set out to determine the identity of the anonymous emailer.

That is the first disturbing fact: it appears that the FBI not only devoted substantial resources, but also engaged in highly invasive surveillance, for no reason other than to do a personal favor for a friend of one of its agents, to find out who was very mildly harassing her by email.

Think about that. If an FBI agent can go digging through private emails over a friend's complaint about nasty-grams, doesn't that suggest that such intrusive snooping is pretty much old hat to the feds?

Greenwald points out that the FBI's digging into Paula Broadwell's nasty-grams not only took them into her email account and revealed her relationship with David Petraeus; it then revealed Jill Kelley's correspondence with General John Allen, including a truly awe-inspiring data-dump of emails between the two. Continues Greenwald:

So not only did the FBI - again, all without any real evidence of a crime - trace the locations and identity of Broadwell and Petreaus, and read through Broadwell's emails (and possibly Petraeus'), but they also got their hands on and read through 20,000-30,000 pages of emails between Gen. Allen and Kelley.

This is a surveillance state run amok. It also highlights how any remnants of internet anonymity have been all but obliterated by the union between the state and technology companies.

Online email services are especially vulnerable, with companies like Google and Yahoo essentially rolling over for the feds. As the Associated Press reported:

The downfall of CIA Director David Petraeus demonstrates how easy it is for federal law enforcement agents to examine emails and computer records if they believe a crime was committed. With subpoenas and warrants, the FBI and other investigating agencies routinely gain access to electronic inboxes and information about email accounts offered by Google, Yahoo and other Internet providers.

In fact, older emails — those six months old or older — don't require a warrant at all. Prosecutors can grab them on their own authority. Many companies will cough up detailed information without a formal warrant, anyway. "Google, which operates the widely used Gmail service, complied with more than 90 percent of the nearly 12,300 requests it received in 2011 from the U.S. government for data about its users, according to figures from the company."

Some email providers have been so eager to comply that they actually surrender more information than the FBI requests — and more than it is legally authorized to seek. One such high-profile incident occurred in 2006.

A technical glitch gave the F.B.I. access to the e-mail messages from an entire computer network — perhaps hundreds of accounts or more — instead of simply the lone e-mail address that was approved by a secret intelligence court as part of a national security investigation, according to an internal report of the 2006 episode.

F.B.I. officials blamed an “apparent miscommunication” with the unnamed Internet provider, which mistakenly turned over all the e-mail from a small e-mail domain for which it served as host. The records were ultimately destroyed, officials said.

So remember ... Your online privacy isn't so private.


11 Nov 05:47

Groupoid cardinality

by Qiaochu Yuan

Suitably nice groupoids have a numerical invariant attached to them called groupoid cardinality. Groupoid cardinality is closely related to Euler characteristic and can be thought of as providing a notion of integration on groupoids.

There are various situations in mathematics where computing the size of a set is difficult but where that set has a natural groupoid structure and computing its groupoid cardinality turns out to be easier and give a nicer answer. In such situations the groupoid cardinality is also known as “mass,” e.g. in the Smith-Minkowski-Siegel mass formula for lattices. There are related situations in mathematics where one needs to describe a reasonable probability distribution on some class of objects and groupoid cardinality turns out to give the correct such distribution, e.g. the Cohen-Lenstra heuristics for class groups. We will not discuss these situations, but they should be strong evidence that groupoid cardinality is a natural invariant to consider.

Axiomatics

For convenience, in this section we will restrict to essentially finite groupoids, namely those groupoids equivalent to groupoids with finitely many objects and morphisms.

Associated to any essentially finite groupoid X is a rational number, its groupoid cardinality \chi(X), which is uniquely determined by the following four properties, analogous to the properties uniquely specifying Euler characteristic:

  1. Cardinality: \chi(\text{pt}) = 1, where \text{pt} is the groupoid with one object and one morphism.
  2. Homotopy invariance: If X \sim Y (X is equivalent to Y), then \chi(X) = \chi(Y).
  3. Gluing: \chi(X \sqcup Y) = \chi(X) + \chi(Y).
  4. Covering: If F : X \to Y is an n-sheeted covering map, then \chi(X) = n \chi(Y).

A covering map of groupoids is a functor F : X \to Y which is surjective on objects and which satisfies the unique path lifting property: if p : y_1 \to y_2 is a morphism in Y and x_1 is an object in X such that F(x_1) = y_1, then there exists a unique morphism p' : x_1 \to x_2 in X such that F(p') = p. This axiomatizes the path lifting property satisfied by a covering map of topological spaces. A covering map is n-sheeted if the preimage of every object in Y consists of n objects in X.

The homotopy invariance and gluing axioms imply that groupoid cardinality is completely determined by how it behaves on one-object groupoids BG, where G is a finite group (since we are assuming essential finiteness). Associated to any such groupoid is a canonical |G|-sheeted cover

\displaystyle EG \to BG

where EG is the action groupoid for the action of G on itself (the objects are the elements of G and there is a unique morphism g \to h between any pair of objects). This covering map sends the morphism g \to h to the element h g^{-1} of G. The notation EG is by strong analogy with the theory of classifying spaces.

Since EG is equivalent to a point, \chi(EG) = 1 by the cardinality axiom, and the covering axiom then implies that \chi(BG) = \frac{1}{|G|}. In conclusion, we find that if X is an essentially finite groupoid then, writing the skeleton of X as

\displaystyle \bigsqcup_{x \in \pi_0(X)} B\text{Aut}(x)

we have

\displaystyle \chi(X) = \sum_{x \in \pi_0(X)} \frac{1}{|\text{Aut}(x)|}.

In words, the groupoid cardinality of X is a weighted sum over the isomorphism classes of objects in X, where an object is weighted by the size of its automorphism group. Intuitively speaking, we can think of the objects of X as being “cut up” by their automorphism groups into fractional points.

Groupoid cardinality has other properties besides the above that make it a natural measure of the size of a groupoid.

Proposition: Let X, Y be essentially finite groupoids. Then their product X \times Y is also essentially finite, and \chi(X \times Y) = \chi(X) \times \chi(Y).

Proof. A groupoid is essentially finite if and only if it has finitely many isomorphism classes and the objects in each isomorphism class have finitely many automorphisms. This condition is preserved under finite products; moreover, if

\displaystyle X \sim \bigsqcup_{x \in \pi_0(X)} B \text{Aut}(x)

and

\displaystyle Y \sim \bigsqcup_{y \in \pi_0(Y)} B \text{Aut}(y)

then

\displaystyle X \times Y \sim \bigsqcup_{(x, y) \in (\pi_0(X) \times \pi_0(Y))} B(\text{Aut}(x) \times \text{Aut}(y))

which gives the desired result. \Box

Alternatively, one could show that \frac{\chi(- \times Y)}{\chi(Y)} satisfied all of the axioms above.

Proposition: Let S be a finite set and G be a finite group acting on S. Then the groupoid cardinality of the action groupoid or weak quotient S//G is \chi(S//G) = \frac{|S|}{|G|}.

Note that this is badly false for the set-theoretic quotient S/G, a point which trips up many beginners in combinatorics.

The idea of the proof is that we would like to apply the covering axiom to the natural map S \to S//G (thinking of S as a discrete groupoid), except that this map isn’t a covering map unless the action of G is free. However, it can be replaced by a covering map up to equivalence (a kind of fibrant replacement) essentially using the Borel construction.

Proof. Instead of considering S, consider the equivalent groupoid S \times EG, which consists of pairs (s, g) where s \in S, g \in G, and where there is a unique morphism (s, g) \to (s, h) for every g, h \in G. Since G acts on both S and EG, it acts on this product, and so we can consider the action groupoid (S \times EG)//G and the corresponding map

\displaystyle S \times EG \to (S \times EG)//G.

Since G acts freely on S \times EG, this map is a |G|-sheeted covering map. Moreover, S \times EG \sim S and (S \times EG)//G \sim S//G. We can now apply the covering axiom, and the conclusion follows. \Box

For a more pedestrian proof, observe that it suffices by the gluing axiom to prove the statement in the case that the action of G on S is transitive, where it reduces to the orbit-stabilizer theorem.

Digression: random finite sets

The definition of groupoid cardinality can be extended to tame groupoids, namely those groupoids X such that the sum

\displaystyle \sum_{x \in \pi_0(X)} \frac{1}{|\text{Aut}(x)|}

converges. For any such groupoid, there is a natural probability measure on \pi_0(X) given by the condition that a given isomorphism class x \in \pi_0(X) occurs with probability

\displaystyle \frac{1}{\chi(X)} \left( \frac{1}{|\text{Aut}(x)|} \right).

For example, if X = \text{core}(\text{FinSet}) is the groupoid of finite sets and bijections, then

\displaystyle \chi(X) = \sum_{n \ge 0} \frac{1}{n!} = e

and the finite set of cardinality n occurs with probability \frac{1}{e n!}. In other words, “size of a random finite set” is Poisson with parameter \lambda = 1. It is unclear to me what the significance of this observation is, if any.

More generally, let s be a finite set and consider the groupoid of s-colored finite sets. This is the groupoid whose objects are finite sets x equipped with a map x \to s (assigning to each element of x its color) and whose morphisms are bijections x_1 \to x_2 compatible with colors. The cardinality of this groupoid may be computed in two ways. On the one hand, there are s^n isomorphism types of objects where |x| = n, and the groupoid consisting these isomorphism types is equivalent to the action groupoid of S_n acting on the set of all functions from an n-element set to s, hence the groupoid cardinality is

\displaystyle \sum_{n \ge 0} \frac{|s|^n}{n!}.

On the other hand, the groupoid of s-colored finite sets is equivalent to the product of |s| copies of the groupoid of finite sets; the equivalence is given by sending an s-colored finite set to the finite sets given by the elements of each color. It is not hard to show that for tame groupoids we have \chi(X \times Y) = \chi(X) \times \chi(Y), hence the groupoid cardinality is

\displaystyle \left( \sum_{n \ge 0} \frac{1}{n!} \right)^{|s|}.

Hence “size of a random s-colored finite set” is Poisson with parameter \lambda = |s|, and along the way to seeing this we have shown that two ways of defining e^{|s|} give the same answer (and also implicitly given a combinatorial proof that e^{|s| + |t|} = e^{|s|} e^{|t|}).

There is much more to say about these kinds of arguments, much of which has been said by John Baez at some point, but I don’t know a place where all of the relevant links have been collected. One place to start and work backwards from is week300.

Groupoid cardinality and Euler characteristic

The axiomatic definition of groupoid cardinality suggests that it ought to behave like Euler characteristic, except that the Euler characteristic of familiar spaces are integers and groupoid cardinality is not an integer. However, there is a nice sense in which the Euler characteristic of BG ought to be \frac{1}{|G|}.

BG is a groupoid model of a classifying space of G, also denoted BG, which for discrete groups has two equivalent definitions. It is the unique (up to homotopy) connected space such that \pi_1(BG) = G and such that all higher homotopy groups are trivial; in other words, it is the Eilenberg-MacLane space K(G, 1). Such spaces are also known as aspherical spaces.

The classifying space BG is also the space which represents, in a suitable homotopy category, the functor sending a topological space to its set of principal G-bundles. When G is a discrete group, this is the same thing as a G-cover, but the definition in terms of bundles also generalizes to topological groups.

Example. B\mathbb{Z} is the circle S^1.

Example. More generally, a nice connected space X is a BG for G = \pi_1(X) if and only if its universal cover is contractible; in particular any hyperbolic manifold has this property.

Example. B\mathbb{Z}/2\mathbb{Z} is infinite real projective space \mathbb{RP}^{\infty}.

The sense in which \chi(BG) ought to be \frac{1}{|G|} for G finite is the following. Recall that if X is, say, a finite CW complex, we should have

\displaystyle \chi(X) = \sum_{i \ge 0} (-1)^i c_i

where c_i is the number of i-cells of X. There is a distinguished model of BG (the space) having a cell decomposition in which c_i = (|G| - 1)^i, and thus we ought to have

\displaystyle \chi(BG) = \sum_{i \ge 0} (-1)^i (|G| - 1)^i = \frac{1}{1 + (|G| - 1)} = \frac{1}{|G|}

by summing a divergent geometric series! I learned this from MO. This can be seen more explicitly for \mathbb{RP}^{\infty}, for example, which has a single cell in each dimension and therefore whose Euler characteristic ought to be Grandi’s series

\displaystyle \chi(\mathbb{RP}^{\infty}) = 1 - 1 + 1 \mp ... = \frac{1}{2}.


07 Nov 00:20

Introduction to string diagrams

by Qiaochu Yuan

Today I would like to introduce a diagrammatic notation for dealing with tensor products and multilinear map. The basic idea for this notation appears to be due to Penrose. It has the advantage of both being widely applicable and easier and more intuitive to work with; roughly speaking, computations are performed by topological manipulations on diagrams, revealing the natural notation to use here is 2-dimensional (living in a plane) rather than 1-dimensional (living on a line).

For the sake of accessibility we will restrict our attention to vector spaces. There are category-theoretic things happening in this post but we will not point them out explicitly. We assume familiarity with the notion of tensor product of vector spaces but not much else.

Below the composition of a map f : a \to b with a map g : b \to c will be denoted f \circ g : a \to c (rather than the more typical g \circ f). This will make it easier to translate between diagrams and non-diagrams. All diagrams were drawn in Paper.

String diagrams

String diagrams for finite-dimensional vector spaces work as follows. To start with, a linear map f : U \to V is represented by a box labeled f with one input string labeled U and one output string labeled V. Composition of linear maps f : U \to V and g : V \to W is given by pairing input and output wires with matching labels.

The tensor product of two linear maps f : U_1 \to V_1, g : U_2 \to V_2 is a map f \otimes g : U_1 \otimes U_2 \to V_1 \otimes V_2 represented graphically by stacking boxes vertically. Note that f \otimes g has two input wires and two output wires.

The 1-dimensional vector space 1 is not represented by a wire at all, to reflect the fact that it is an identity for tensor product in the sense that there is a natural isomorphism V \otimes 1 \cong V. Thus a vector in a vector space is a morphism v : 1 \to V, which is just a box with no input wire and one output wire, and a linear functional or covector is a morphism f : V \to 1, which is just a box with no output wire and one input wire.

The identity map \text{id}_V : V \to V is represented by a wire with no attached box, to reflect the fact that it is an identity for composition, and the identity map \text{id}_1 : 1 \to 1 is not represented by anything at all.

In general, an arbitrary linear map

\displaystyle f : U_1 \otimes ... \otimes U_n \to V_1 \otimes ... \otimes V_m

(obtained for example by taking the tensor product of various other maps) is represented by a box labeled f with n input wires labeled U_1, ... U_n and m output wires labeled V_1, ... V_m.

Such maps admit a generalized notion of composition given by pairing only some input and output wires rather than all of them (defined formally by composition after tensoring with a suitable collection of identity morphisms). For example, if m : A \otimes A \to A is a bilinear map on a vector space A, the following is the statement that m is associative:

In 1-dimensional notation, this reads

\displaystyle (m \otimes \text{id}_A) \circ m = (\text{id}_A \otimes m) \circ m.

Implicit in our use of string diagram notation is the interchange law, which asserts that the following diagram is a well-defined map U_1 \otimes U_2 \to W_1 \otimes W_2 (in the sense that the two ways of evaluating it using tensor products and compositions produce the same result):

In 1-dimensional notation, this reads

\displaystyle (f_1 \otimes f_2) \circ (g_1 \otimes g_2) = (f_1 \circ g_1) \otimes (f_2 \circ g_2).

The interchange law should be thought of as a 2-dimensional version of associativity. It allows us to “drop generalized parentheses” in the sense that we do not have to specify what order we tensor and compose a family of maps as above.

Symmetry

The tensor product is commutative in a suitable sense, so we should be able to freely change the order of input and output wires. We can do this formally as follows. For any pair U, V of vector spaces there is a distinguished symmetry map

\displaystyle \gamma_{U, V} : U \otimes V \ni u \otimes v \mapsto v \otimes u \in V \otimes U

which is represented by unlabeled crossing wires. The symmetry maps obey various axioms which ensure that they behave like crossing wires ought to. The most important axiom is naturality, which asserts that we can slide boxes along symmetries:

In 1-dimensional notation, this reads

\displaystyle (f \otimes g) \circ \gamma_{U_2, V_2} = \gamma_{U_1, V_1} \circ (g \otimes f).

Note that if we only want to slide one box along a symmetry we can let the other one be an identity.

Strictly speaking, naturality also applies to morphisms drawn with more than one input or output wire, so can look more complicated than the above.

Another axiom obeyed by the symmetry asserts that applying the symmetry twice gives the identity. Topologically it is described by pulling two wires apart. Looking ahead to future posts, we will call this axiom Reidemeister II:

In 1-dimensional notation, this reads \gamma_{U, V} \circ \gamma_{V, U} = \text{id}_{U \otimes V}.

The third axiom we will discuss is sometimes called the braid relation, but following the pattern of the above naming scheme we will call it Reidemeister III. Topologically it is described by pulling the middle wire across a crossing:

In 1-dimensional notation, this reads that

\displaystyle (\text{id}_U \otimes \gamma_{V, W}) \circ (\gamma_{U, W} \otimes \text{id}_V) \circ (\text{id}_W \otimes \gamma_{U, V})

is equal to

\displaystyle (\gamma_{U, V} \otimes \text{id}_W) \circ (\text{id}_V \otimes \gamma_{U, W}) \circ (\gamma_{V, W} \circ \text{id}_U).

In particular, specialized to the case of an n-fold tensor product V^{\otimes n}, Reidemeister II and III are precisely the relations in a well-known presentation of the symmetric group S_n, so that S_n naturally acts on V^{\otimes n}.

The symmetry maps allow string diagrams greater expressive power. For example, if m : A \otimes A \to A is a bilinear map, the following is the statement that m is commutative:


14 Oct 12:29

The ducks are gonna get you [Pharyngula]

by PZ Myers
Nosimpler

Poor girl indeed.

Some poor young girl, deeply miseducated and misled, wrote into a newspaper with a letter trying to denounce homosexuality with a bad historical and biological argument. She’s only 14, and her brain has already been poisoned by the cranks and liars in her own family…it’s very sad. Here’s the letter — I will say, it’s a very creative argument that would be far more entertaining if it weren’t wrong in every particular.

I’ve transcribed it below. I couldn’t help myself, though, and had to, um, annotate it a bit.

Homosexuality, including same sex marriage, is not an enlightened idea [But tolerance and acceptance of diversity are]. The Romans practiced homosexuality [Every culture has had homosexual individuals; they differ only in the degree of suppression. The Romans actually regarded homosexuals as effete and inferior, and used accusations of gayness as expressions of contempt, just like modern middle schoolers]. Surely, after 2000 years, our level of intelligence should have evolved somewhat, so that we can truly pride ourselves of being cleverer than our forebears [Two millennia is actually a short span of time for biological evolution. Also, have you ever heard of the Dark Ages? Progress is not inevitable].

If homosexuality spreads, it can cause human evolution to come to a standstill [Nope. Homosexuals reproduce. Homosexuality refers to behavior and social preferences, not to biological limitations. Also, many heterosexuals choose to not reproduce as well, and it does not stop evolution in its tracks — in complex social organisms like ours, there are many ways to contribute to the species that don't involve breeding directly]. It could threaten the human position on the evolutionary ladder [There is no evolutionary "ladder". You have some serious misconceptions about biology, young lady!], and say, ducks, could take over the world [Evolution is not about taking over the world. There is no pinnacle. Every species has a different niche, not a different spot in a hierarchy of dominance]. Ducks always nest in pairs [This is called the naturalistic fallacy. You cannot draw conclusions from how one species behaves and declare that it justifies one specific kind of behavior in another species. I could point to gorillas, and announce that we should live in polygamous harems; I could point to bonobos and say that public homosexual acts ought to be accepted as a matter of course, and that we ought to have casual sex as often as we say hello. If you'd like, I could give you a long list of very kinky sexual behaviors practiced by various species on the planet; shall we decide that because ducks rape, so should we, lest we fall behind evolutionarily?] and if we allow same-sex marriage, then the ducks will have evolved further than we have [Ducks are just as "evolved" as we are, and we're not more evolved than any other species on the planet. Evolution is about branching trees, not climbing ladders]. We will be in danger of all being equal, with ducks more equal than us [That makes no sense].

We should learn from history and not be stuck with copying ancient behavior [Are you, by any chance, a follower of Jesus or Mohammed? Because you know, those faiths are all about imposing ancient rules for behavior on modern society]. The government has no right to bring us back to the stone age [But the Middle Ages are OK, I suppose?]. I don’t want my children to have to compete with ducks [Wait. I'm trying to puzzle this out. Because you think ducks are all heterosexual, and your children will all be heterosexual (brace yourself, you might get a few surprises in 10 or 20 years there), and a policy of tolerance will turn every other human being homosexual, you're afraid your kids will be competing for mates with ducks? Or is it that duck heterosexuality is the only criterion that makes them acceptable for positions of power, so years from now, your children will find themselves in a workplace dominated by duck bosses, who have overcome the handicap of lack of manipulatory appendages and very small brains to be in charge of everything? I don't get it]. I want them to evolve further than I have [But you don't believe in evolution!]. Any self-respecting human would aim for that, too. [Are you aware that the Abrahamic faiths all preach that humanity is in a state of ineluctable decay since the Fall and that human sin corrupts us? I don't think any self-respecting human should be a Christian or a Jew or Muslim, for the same reason]

None of this really bears any weight for be, because I do not believe in evolution [You don't understand it, either]. However, the powers that be believe in evolution, and have made many decisions based on it. They should be consistent: if you believe in evolution, then you can’t be in favour of homosexuality [If you accept evolution, then you recognize that there are diverse successful sexual strategies in the world, and you also have a deeper appreciation of the complexity of biology, so no, you should be much more accepting of reality], or the ducks will get you in the end [You can live your life in fear of ducks, or you can love your fellow human beings and encourage more love in the world. Your choice].

Jasmin H, aged 14 [You have time to grow up!]
Homeschooled [Obviously], Scargill

11 Oct 22:43

The classical mechanics of non-conservative systems. (arXiv:1210.2745v2 [gr-qc] UPDATED)

by Chad R. Galley
Nosimpler

We may be out of a job guys.

Hamilton's principle of stationary action lies at the foundation of theoretical physics and is applied in many other disciplines from pure mathematics to economics. Despite its utility, Hamilton's principle has a subtle pitfall that often goes unnoticed in physics: it is formulated as a boundary value problem in time but is used to derive equations of motion that are solved with initial data. This subtlety can have undesirable effects. I present a formulation of Hamilton's principle that is compatible with initial value problems. Remarkably, this leads to a natural formulation for the Lagrangian and Hamiltonian dynamics of generic non-conservative systems, thereby filling a long-standing gap in classical mechanics. Thus dissipative effects, for example, can be studied with new tools that may have application in a variety of disciplines. The new formalism is demonstrated by two examples of non-conservative systems: an object moving in a fluid with viscous drag forces and a harmonic oscillator coupled to a dissipative environment.

10 Oct 22:54

Rocking the foundations of molecular genetics [Commentary]

by Mattick, J. S.
In PNAS, Nelson et al. present intriguing evidence that challenges the fundamental tenets of genetics (1). It has long been assumed that the inherited contribution to phenotype is embedded in DNA sequence variations in, and interactions between, the genes endogenous to the organism, i.e., alleles derived from parents with some...
10 Oct 22:47

The Measurement That Would Reveal The Universe As A Computer Simulation

Nosimpler

Hey this one might actually make sense kinda.

If the cosmos is a numerical simulation, there ought to be clues in the spectrum of high energy cosmic rays, say theorists

One of modern physics' most cherished ideas is quantum chromodynamics, the theory that describes the strong nuclear force, how it binds quarks and gluons into protons and neutrons, how these form nuclei that themselves interact. This is the universe at its most fundamental. 

So an interesting pursuit is to simulate quantum chromodynamics on a computer to see what kind of complexity arises. The promise is that simulating physics on such a fundamental level is more or less equivalent to simulating the universe itself.  

There are one or two challenges of course. The physics is mind-bogglingly complex and operates on a vanishingly small scale. So even using the world's most powerful supercomputers, physicists have only managed to simulate tiny corners of the cosmos just a few femtometers across. (A femtometer is 10^-15 metres.) 

That may not sound like much but the significant point is that the simulation is essentially indistinguishable from the real thing (at least as far as we understand it).  

It's not hard to imagine that Moore's Law-type progress will allow physicists to simulate significantly larger regions of space. A region just a few micrometres across could encapsulate the entire workings of a human cell. 

Again, the behaviour of this human cell would be indistinguishable from the real thing.

It's this kind of thinking that forces physicists to consider the possibility that our entire cosmos could be running on a vastly powerful computer. If so, is there any way we could ever know?   

Today, we get an answer of sorts from Silas Beane, at the University of Bonn in Germany, and a few pals.  They say there is a way to see evidence that we are being simulated, at least in certain scenarios.

First, some background. The problem with all simulations is that the laws of physics, which appear continuous, have to be superimposed onto a discrete three dimensional lattice which advances in steps of time. 

The question that Beane and co ask is whether the lattice spacing imposes any kind of limitation on the physical processes we see in the universe. They examine, in particular, high energy processes, which probe smaller regions of space as they get more energetic 

What they find is interesting. They say that the lattice spacing imposes a fundamental limit on the energy that particles can have. That's because nothing can exist that is smaller than the lattice itself. 

So if our cosmos is merely a simulation, there ought to be a cut off in the spectrum of high energy particles.

It turns out there is exactly this kind of cut off in the energy of cosmic ray particles,  a limit known as the Greisen–Zatsepin–Kuzmin or GZK cut off. 

This cut-off has been well studied and comes about because high energy particles interact with the cosmic microwave background and so lose energy as they travel  long distances. 

But Beane and co calculate that the lattice spacing imposes some additional features on the spectrum. "The most striking feature...is that the angular distribution of the highest energy components would exhibit cubic symmetry in the rest frame of the lattice, deviating significantly from isotropy," they say.

In other words, the cosmic rays would travel preferentially along the axes of the lattice, so we wouldn't see them equally in all directions. 

That's a measurement we could do now with current technology. Finding the effect would be equivalent to being able to to 'see' the orientation of lattice on which our universe is simulated.

That's cool, mind-blowing even. But the calculations by Beane and co are not without some important caveats. One problem is that the computer lattice may be constructed in an entirely different way to the one envisaged by these guys.  

Another is that this effect is only measurable if the lattice cut off is the same as the GZK cut off. This occurs when the lattice spacing is about 10^-12 femtometers. If the spacing is significantly smaller than that, we'll see nothing.

Nevertheless, it's surely worth looking for, if only to rule out the possibility that we're part of a simulation of this particular kind but secretly in the hope that we'll find good evidence of our robotic overlords once and for all.

Ref: arxiv.org/abs/1210.1847: Constraints on the Universe as a Numerical Simulation



08 Oct 18:04

“Observer Space”: Cartan Geometry and Lifting General Relativity

by Jeffrey Morton

This entry is a by-special-request blog, which Derek Wise invited me to write for the blog associated with the International Loop Quantum Gravity Seminar, and it will appear over there as well.  The ILQGS is a long-running regular seminar which runs as a teleconference, with people joining in from various countries, on various topics which are more or less closely related to Loop Quantum Gravity and the interests of people who work on it.  The custom is that when someone gives a talk, someone else writes up a description of the talk for the ILQGS blog, and Derek invited me to write up a description of his talk.  The audio file of the talk itself is available in .aiff and .wav formats, and the slides are here.

The talk that Derek gave was based on a project of his and Steffen Gielen’s, which has taken written form in a few papers (two shorter ones, “Spontaneously broken Lorentz symmetry for Hamiltonian gravity“, “Linking Covariant and Canonical General Relativity via Local Observers“, and a new, longer one called “Lifting General Relativity to Observer Space“).

The key idea behind this project is the notion of “observer space”, which is exactly what it sounds like: a space of all observers in a given universe.  This is easiest to picture when one has a spacetime – a manifold with a Lorentzian metric, (M,g) – to begin with.  Then an observer can be specified by choosing a particular point (x_0,x_1,x_2,x_3) = \mathbf{x} in spacetime, as well as a unit future-directed timelike vector v.  This vector is a tangent to the observer’s worldline at \mathbf{x}.  The observer space is therefore a bundle over M, the “future unit tangent bundle”.  However, using the notion of a “Cartan geometry”, one can give a general definition of observer space which makes sense even when there is no underlying (M,g).

The result is a surprising, relatively new physical intuition is that “spacetime” is a local and observer-dependent notion, which in some special cases can be extended so that all observers see the same spacetime.  This is somewhat related to the relativity of locality, which I’ve blogged about previously.  Geometrically, it is similar to the fact that a slicing of spacetime into space and time is not unique, and not respected by the full symmetries of the theory of Relativity, even for flat spacetime (much less for the case of General Relativity).  Similarly, we will see a notion of “observer space”, which can sometimes be turned into a bundle over an objective spacetime M, but not in all cases.

So, how is this described mathematically?  In particular, what did I mean up there by saying that spacetime becomes observer-dependent?

Cartan Geometry

The answer uses Cartan geometry, which is a framework for differential geometry that is slightly broader than what is commonly used in physics.  Roughly, one can say “Cartan geometry is to Klein geometry as Riemannian geometry is to Euclidean geometry”.  The more familiar direction of generalization here is the fact that, like Riemannian geometry, Cartan is concerned with manifolds which have local models in terms of simple, “flat” geometries, but which have curvature, and fail to be homogeneous.  First let’s remember how Klein geometry works.

Klein’s Erlangen Program, carried out in the mid-19th-century, systematically brought abstract algebra, and specifically the theory of Lie groups, into geometry, by placing the idea of symmetry in the leading role.  It describes “homogeneous spaces”, which are geometries in which every point is indistinguishable from every other point.  This is expressed by the existence of a transitive action of some Lie group G of all symmetries on an underlying space.  Any given point x will be fixed by some symmetries, and not others, so one also has a subgroup H = Stab(x) \subset G.  This is the “stabilizer subgroup”, consisting of all symmetries which fix x.  That the space is homogeneous means that for any two points x,y, the subgroups Stab(x) and Stab(y) are conjugate (by a symmetry taking x to y).  Then the homogeneous space, or Klein geometry, associated to (G,H) is, up to isomorphism, just the same as the quotient space G/H of the obvious action of H on G.

The advantage of this program is that it has a great many examples, but the most relevant ones for now are:

  • n-dimensional Euclidean space. the Euclidean group ISO(n) = SO(n) \ltimes \mathbb{R}^n is precisely the group of transformations that leave the data of Euclidean geometry, lengths and angles, invariant.  It acts transitively on \mathbb{R}^n.  Any point will be fixed by the group of rotations centred at that point, which is a subgroup of ISO(n) isomorphic to SO(n).  Klein’s insight is to reverse this: we may define Euclidean space by R^n \cong ISO(n)/SO(n).
  • n-dimensional Minkowski space.  Similarly, we can define this space to be ISO(n-1,1)/SO(n-1,1).  The Euclidean group has been replaced by the Poincaré group, and rotations by the Lorentz group (of rotations and boosts), but otherwise the situation is essentially the same.
  • de Sitter space.  As a Klein geometry, this is the quotient SO(4,1)/SO(3,1).  That is, the stabilizer of any point is the Lorentz group – so things look locally rather similar to Minkowski space around any given point.  But the global symmetries of de Sitter space are different.  Even more, it looks like Minkowski space locally in the sense that the Lie algebras give representations so(4,1)/so(3,1) and iso(3,1)/so(3,1) are identical, seen as representations of SO(3,1).  It’s natural to identify them with the tangent space at a point.  de Sitter space as a whole is easiest to visualize as a 4D hyperboloid in \mathbb{R}^5.  This is supposed to be seen as a local model of spacetime in a theory in which there is a cosmological constant that gives empty space a constant negative curvature.
  • anti-de Sitter space. This is similar, but now the quotient is SO(3,2)/SO(3,1) – in fact, this whole theory goes through for any of the last three examples: Minkowski; de Sitter; and anti-de Sitter, each of which acts as a “local model” for spacetime in General Relativity with the cosmological constant, respectively: zero; positive; and negative.

Now, what does it mean to say that a Cartan geometry has a local model?  Well, just as a Lorentzian or Riemannian manifold is “locally modelled” by Minkowski or Euclidean space, a Cartan geometry is locally modelled by some Klein geometry.  This is best described in terms of a connection on a principal G-bundle, and the associated G/H-bundle, over some manifold M.  The crucial bundle in a Riemannian or Lorenztian geometry is the frame bundle: the fibre over each point consists of all the ways to isometrically embed a standard Euclidean or Minkowski space into the tangent space.  A connection on this bundle specifies how this embedding should transform as one moves along a path.  It’s determined by a 1-form on M, valued in the Lie algebra of G.

Given a parametrized path, one can apply this form to the tangent vector at each point, and get a Lie algebra-valued answer.  Integrating along the path, we get a path in the Lie group G (which is independent of the parametrization).  This is called a “development” of the path, and by applying the G-values to the model space G/H, we see that the connection tells us how to move through a copy of G/H as we move along the path.  The image this suggests is of “rolling without slipping” – think of the case where the model space is a sphere.  The connection describes how the model space “rolls” over the surface of the manifold M.  Curvature of the connection measures the failure to commute of the processes of rolling in two different directions.  A connection with zero curvature describes a space which (locally at least) looks exactly like the model space: picture a sphere rolling against its mirror image.  Transporting the sphere-shaped fibre around any closed curve always brings it back to its starting position. Now, curvature is defined in terms of transports of these Klein-geometry fibres.  If curvature is measured by the development of curves, we can think of each homogeneous space as a flat Cartan geometry with itself as a local model.

This idea, that the curvature of a manifold depends on the model geometry being used to measure it, shows up in the way we apply this geometry to physics.

Gravity and Cartan Geometry

MacDowell-Mansouri gravity can be understood as a theory in which General Relativity is modelled by a Cartan geometry.  Of course, a standard way of presenting GR is in terms of the geometry of a Lorentzian manifold.  In the Palatini formalism, the basic fields are a connection A and a vierbein (coframe field) called e, with dynamics encoded in the Palatini action, which is the integral over M of R[\omega] \wedge e \wedge e, where R is the curvature 2-form for \omega.

This can be derived from a Cartan geometry, whose model geometry is de Sitter space SO(4,1)/SO(3,1).   Then MacDowell-Mansouri gravity gets \omega and e by splitting the Lie algebra as so(4,1) = so(3,1) \oplus \mathbb{R^4}.  This “breaks the full symmetry” at each point.  Then one has a fairly natural action on the so(4,1)-connection:

\int_M tr(F_h \wedge \star F_h)

Here, F_h is the so(3,1) part of the curvature of the big connection.  The splitting of the connection means that F_h = R + e \wedge e, and the action above is rewritten, up to a normalization, as the Palatini action for General Relativity (plus a topological term, which has no effect on the equations of motion we get from the action).  So General Relativity can be written as the theory of a Cartan geometry modelled on de Sitter space.

The cosmological constant in GR shows up because a “flat” connection for a Cartan geometry based on de Sitter space will look (if measured by Minkowski space) as if it has constant curvature which is exactly that of the model Klein geometry.  The way to think of this is to take the fibre bundle of homogeneous model spaces as a replacement for the tangent bundle to the manifold.  The fibre at each point describes the local appearance of spacetime.  If empty spacetime is flat, this local model is Minkowski space, ISO(3,1)/SO(3,1), and one can really speak of tangent “vectors”.  The tangent homogeneous space is not linear.  In these first cases, the fibres are not vector spaces, precisely because the large group of symmetries doesn’t contain a group of translations, but they are Klein geometries constructed in just the same way as Minkowski space. Thus, the local description of the connection in terms of Lie(G)-valued forms can be treated in the same way, regardless of which Klein geometry G/H occurs in the fibres.  In particular, General Relativity, formulated in terms of Cartan geometry, always says that, in the absence of matter, the geometry of space is flat, and the cosmological constant is included naturally by the choice of which Klein geometry is the local model of spacetime.

Observer Space

The idea in defining an observer space is to combine two symmetry reductions into one.  The reduction from SO(4,1) to SO(3,1) gives de Sitter space, SO(4,1)/SO(3,1) as a model Klein geometry, which reflects the “symmetry breaking” that happens when choosing one particular point in spacetime, or event.  Then, the reduction of SO(3,1) to SO(3) similarly reflects the symmetry breaking that occurs when one chooses a specific time direction (a future-directed unit timelike vector).  These are the tangent vectors to the worldline of an observer at the chosen point, so SO(3,1)/SO(3) the model Klein geometry, is the space of such possible observers.  The stabilizer subgroup for a point in this space consists of just the rotations of space around the corresponding observer – the boosts in SO(3,1) translate between observers.  So locally, choosing an observer amounts to a splitting of the model spacetime at the point into a product of space and time. If we combine both reductions at once, we get the 7-dimensional Klein geometry SO(4,1)/SO(3).  This is just the future unit tangent bundle of de Sitter space, which we think of as a homogeneous model for the “space of observers”

A general observer space O, however, is just a Cartan geometry modelled on SO(4,1)/SO(3).  This is a 7-dimensional manifold, equipped with the structure of a Cartan geometry.  One class of examples are exactly the future unit tangent bundles to 4-dimensional Lorentzian spacetimes.  In these cases, observer space is naturally a contact manifold: that is, it’s an odd-dimensional manifold equipped with a 1-form \alpha, the contact form, which is such that the top-dimensional form \alpha \wedge d \alpha \wedge \dots \wedge d \alpha is nowhere zero.  This is the odd-dimensional analog of a symplectic manifold.  Contact manifolds are, intuitively, configuration spaces of systems which involve “rolling without slipping” – for instance, a sphere rolling on a plane.  In this case, it’s better to think of the local space of observers which “rolls without slipping” on a spacetime manifold M.

Now, Minkowski space has a slicing into space and time – in fact, one for each observer, who defines the time direction, but the time coordinate does not transform in any meaningful way under the symmetries of the theory, and different observers will choose different ones.  In just the same way, the homogeneous model of observer space can naturally be written as a bundle SO(4,1)/SO(3) \rightarrow SO(4,1)/SO(3,1).  But a general observer space O may or may not be a bundle over an ordinary spacetime manifold, O \rightarrow M.  Every Cartan geometry M gives rise to an observer space O as the bundle of future-directed timelike vectors, but not every Cartan geometry O is of this form, in any natural way. Indeed, without a further condition, we can’t even reconstruct observer space as such a bundle in an open neighborhood of a given observer.

This may be intuitively surprising: it gives a perfectly concrete geometric model in which “spacetime” is relative and observer-dependent, and perhaps only locally meaningful, in just the same way as the distinction between “space” and “time” in General Relativity. It may be impossible, that is, to determine objectively whether two observers are located at the same base event or not. This is a kind of “Relativity of Locality” which is geometrically much like the by-now more familiar Relativity of Simultaneity. Each observer will reach certain conclusions as to which observers share the same base event, but different observers may not agree.  The coincident observers according to a given observer are those reached by a good class of geodesics in O moving only in directions that observer sees as boosts.

When one can reconstruct O \rightarrow M, two observers will agree whether or not they are coincident.  This extra condition which makes this possible is an integrability constraint on the action of the Lie algebra H (in our main example, H = SO(3,1)) on the observer space O.  In this case, the fibres of the bundle are the orbits of this action, and we have the familiar world of Relativity, where simultaneity may be relative, but locality is absolute.

Lifting Gravity to Observer Space

Apart from describing this model of relative spacetime, another motivation for describing observer space is that one can formulate canonical (Hamiltonian) GR locally near each point in such an observer space.  The goal is to make a link between covariant and canonical quantization of gravity.  Covariant quantization treats the geometry of spacetime all at once, by means of a Lagrangian action functional.  This is mathematically appealing, since it respects the symmetry of General Relativity, namely its diffeomorphism-invariance.  On the other hand, it is remote from the canonical (Hamiltonian) approach to quantization of physical systems, in which the concept of time is fundamental. In the canonical approach, one gets a Hilbert space by quantizing the space of states of a system at a given point in time, and the Hamiltonian for the theory describes its evolution.  This is problematic for diffeomorphism-, or even Lorentz-invariance, since coordinate time depends on a choice of observer.  The point of observer space is that we consider all these choices at once.  Describing GR in O is both covariant, and based on (local) choices of time direction.

This is easiest to describe in the case of a bundle O \rightarrow M.  Then a “field of observers” to be a section of the bundle: a choice, at each base event in M, of an observer based at that event.  A field of observers may or may not correspond to a particular decomposition of spacetime into space evolving in time, but locally, at each point in O, it always looks like one.  The resulting theory describes the dynamics of space-geometry over time, as seen locally by a given observer.  In this case, a Cartan connection on observer space is described by to a Lie(SO(4,1))-valued form.  This decomposes into four Lie-algebra valued forms, interpreted as infinitesimal transformations of the model observer by: (1) spatial rotations; (2) boosts; (3) spatial translations; (4) time translation.  The four-fold division is based on two distinctions: first, between the base event at which the observer lives, and the choice of observer (i.e. the reduction of SO(4,1) to SO(3,1), which symmetry breaking entails choosing a point); and second, between space and time (i.e. the reduction of SO(3,1) to SO(3), which symmetry breaking entails choosing a time direction).

This splitting, along the same lines as the one in MacDowell-Mansouri gravity described above, suggests that one could lift GR to a theory on an observer space O.  This amount to describing fields on O and an action functional, so that the splitting of the fields gives back the usual fields of GR on spacetime, and the action gives back the usual action.  This part of the project is still under development, but this lifting has been described.  In the case when there is no “objective” spacetime, the result includes some surprising new fields which it’s not clear how to deal with, but when there is an objective spacetime, the resulting theory looks just like GR.


08 Oct 17:50

The Anarchist Fitness Program

by Jesse Walker

James C. Scott, of Seeing Like a State fame, is about to release a new book called Two Cheers for Anarchism. Over at Bleeding Heart Libertarians, Matt Zwolinki quotes a passage from it:

Cartoon by the great Ron Cobb.

One day you will be called upon to break a big law in the name of justice and rationality. Everything will depend on it. You have to be ready. How are you going to prepare for that day when it really matters? You have to stay "in shape" so that when the big day comes you will be ready. What you need is "anarchist calisthenics." Every day or so break some trivial law that makes no sense, even if it’s only jaywalking. Use your own head to judge whether a law is just or reasonable. That way, you'll keep trim; and when the big day comes, you'll be ready.

For the quotation's context, which Godwin's Law aficionados should appreciate, go here.

We have a review of Two Cheers in the works; in the meantime, you can read my review of Seeing Like a State here and Tom Palmer's review of another Scott book here. More Scott cameos in Reason can be found here, here, here, here, and here.


05 Oct 00:10

The NDAA Retroactively "Ass Covers" Some of the More Broadly-Applied Gitmo Detainments Says Lawsuit Plaintiff [Updated/Clarified]

by Lucy Steigerwald

[Note: this piece has been updated with an email clarification from NDAA lawsuit plaintiff Tangerine Bolen. My apologies if I misquoted or misinterpreted anything she said.]*

With 500-some pages of text, the 2012 National Defense Authorization Act (NDAA) covers a lot more than just section 1021(b), but the majority of the debates over the bill involve the very reason the four letters N-D-A-A have become shorthand for fears of government power finally crossing a Rubicon. Whether or not that’s really true, the caginess of the government in respect to who it can indefinitely detain[pdf] is disturbing and demands a clarification that is not being offered. 

Section 1021(b) reads that someone who can be indefinitely detained is:

A person who was a part of or substantially supported al-Qaeda, the Taliban, or associated forces that are engaged in hostilities against the United States or its coalition partners, including any person who has committed a belligerent act or has directly supported such hostilities in aid of such enemy forces.

The government says the controversial bit of the NDAA is nothing new, but seven plaintiffs, including Pentagon Papers leaker Daniel Ellsberg, dissident writer Noam Chomsky, and journalist Chris Hedges, sued in January, arguing that they were under threat. Hedges in particular argued that his First Amendment rights are violated by the NDAA since he has interviewed numerous members of Al-Qaeda and the Taliban, but now fears doing so.

Another plaintiff in Hedges v. Obama is activist Jennifer "Tangerine" Bolen, founder of the pro-whistleblower group RevolutionTruth.org. She worries that her organization's support of WikiLeaks and imprisoned soldier and accused leaker Bradley Manning might also make her or her allies applicable for detainment under the NDAA.

Section 1021(a) of the bill repeats the government's power to go after perpetrators (and those who harbored them, etc.) of the September 11th attacks (put in writing in the joint Authorization for Use of Military Force resolution) but 1021(b) does read an awful lot like it's expanding powers, even if the actual text of the NDAA and Obama administration officials claim it isn't changing anything. (For a good overview of the NDAA up until now, go check out this Young Americans for Liberty blog post.)

Bolen believes part of the subtext to these argument is that the government wants an excuse to go after Julian Assange and Wikileaks."They don't want to go after The New York Times," she says, "They’re willing to cherry-pick who they apply indefinite detention to." But once they can get to Assange, this power will "cascade downward" and then people like Bolen or Hedges could be under threat as well.

The government's initial argument was that the powers granted in provision 1021(b)  were exactly the same as those granted by the AUMF. Yet, argues Bolen, if the AUMF and the NDAA are the same, why is the government so desperate to stop this lawsuit? Why did they appeal less than 24 hours after Judge Katherine Forrest’s permanent block of indefinite detainment on September 13? Why do they claim that block could cause "irreparable harm" to the United States? Well, no harm done for the moment. On Tuesday afternoon, the Second Circuit Court of Appeals ruled, and a three-judge panel stayed Forrest's block until a final decision is reached in December. Until then, or until this hits the Supreme Court, indefinite detainment is back on.

The about the NDAA, says Bolen, is that it's a retroactive "CYA" — cover your ass. "The AUMF powers were so broadly overused for 11 years...this is an attempt [by the Obama administration] to codify powers they never had." The Bush administration's secret prisons and detainment, both at Gitmo and at CIA black sites all over, Bolen says that the AUMF didn't allow any of that, but the NDAA would. NDAA is, says Bolen, an attempt to legalize the past 11 years of the most heated debates of the War on Terror. And Hedges v. Obama is “the latch on Pandora’s box” for proving “this incredibly broad application of the AUMF which was never legal.”

In their Tuesday ruling, the Second Circuit judges wrote [pdf] that it was in "the public interest" to grant the government appeal a stay. Part of their reasoning was that the government finally clarified that the plaintiffs had no reason to fear detainment, meaning that they had no standing to sue in the first place.

When the government initially refused to offer assurances that the plaintiffs could not be detained back in March, this made Judge Forrest more sympathetic to the question of whether the seven individuals indeed had standing to sue. Later, in August, seeing that Forrest was indeed going to block indefinite detainment, the government did try to offer assurances that journalists who were independent were under no threat by offering a clarifying brief. This, according to to Bolen brought up a lot of questions still for the judge. Bolen says Forrest asked, "“Are youtube videos independent? Are you going to form a panel to decide who is independent?" and she was still not satisfied, leading to the Judge's 112-page ruling in which she expressed incredulousness over the government's utter failure to make their case.  [Correction: updated language to reflect better accuracy in the timeline of the case.]

The wording in the government's response brief just does not satisfy any of the plaintiffs and opens up more questions over whether the government may actually be considering keeping an eye on journalists who are not seen as "independent."

Bolen, for her part, thinks that the case will make it to the Supreme Court. But it’s up to her and her fellow-plaintiffs to try to change public opinion to make sure NDAA gets thrown out. As for her opinion on Obama, whose administration is pushing so hard on this, well, it doesn't sound as harsh as you might think. She mentions the near-lies that lead to the Iraq war and says that the government is trained to “out-speak everyone” and that’s what they’re doing again. But “it’s less insidious and less horrific than under Bush. Not to excuse Obama, but he inherited a total nightmare…He can’t suddenly deny himself powers…”

* Bolen's response email to me included these clarifying paragraphs. I have the struck-through quote in my notes, but I am not interested in disputing that, and Bolen has more than made it clear that she opposes Obama on this measure and that he -- in the literal sense -- could have denied himseld the NDAA powers, in spite of roads paved for him by Bush.

Firstly, in reference to the secret prisons, GTMO, etc, I did not say the AUMF did not allow those. What I said was that we believe that the AUMF detention powers were over-broadly applied - subsequently sweeping up innocent people - and definitely people who had nothing to do with 9/11, or are members of Al Qaida or the Taliban - which is the definition of those powers. Those prisons are legal, in fact (perhaps not all of the secret ones - I don't know -.... 

Finally, the quote at the end is not reflective of what I said either. Obama is likely in a position whereby he feels he cannot suddenly deny himself powers on which two administrations have come to rely. He cannot afford a terror attack on his watch, and he is likely convinced he has no choice here. That is quite a bit different than what you quoted me as saying. There is no way I think that Obama can't suddenly deny himself powers - I think he believes that is the case and that he is stuck in a position of political realism that this country does not understand. That does not excuse his willingness to erode civil liberties and undermine human rights just like Bush did - I expected, and expect, him to do better. 


04 Oct 23:58

On Thursday They Were Terrorists; On Friday They Weren't

by Jacob Sullum

As of last Friday, the Mujahedeen-e-Khalq (MEK), a formerly violent Iranian opposition group once allied with Saddam Hussein, no longer appears on the State Department's list of "foreign terrorist organizations" (FTOs). The delisting comes after years of lobbying, assisted by prominent political figures, and legal wrangling, culminating in an appeals court ruling ordering Secretary of State Hillary Clinton to act on the MEK's petition by Monday. Clinton's decision probably means the MEK's supporters, who include former Attorney General Michael Mukasey, former New York Mayor Rudy Giuliani, and former Pennsylvania Gov. Ed Rendell, do not need to worry about being charged with providing "material support" to an FTO by helping the group shed that label. Theoretically, however, they could still be prosecuted if it can be shown that they "coordinated" their advocacy with the MEK prior to Friday.

Over at Popehat, Los Angeles attorney Ken White, a former federal prosecutor, recalls that in 1999 he helped convict a man "for helping terrorists who now aren't terrorists." The defendant helped MEK members "secure legal residence in the United States through various forms of fraud, including fraudulent asylum applications." In addition to immigration fraud, his actions qualified as providing material support to an FTO—possibly the first conviction under that provision, White says. Looking back, he is ambivalent about his role in the case, recognizing the political considerations that determine which groups count as FTOs:

The six people the MEK killed in the 1970s are still dead. They were dead when the State Department designated the MEK as a foreign terrorist organization and they have been dead all the years since and they won't get any less dead when the State Department removes the MEK from its FTO list. The MEK is the organization that once allied with Saddam Hussein; that historical fact hasn't changed, although its political significance has. No — what has changed is the MEK's political power and influence and the attitude of our government towards it.

More generally, White says, the MEK's delisting shows how arbitrary the contours of the War on Terror are:

The scope of the War on Terror — the very identity of the Terror we fight — is a subjective matter in the discretion of the government. The compelling need the government cites to do whatever it wants is itself defined by the government.

The definition of the enemy then determines not only who can be charged with violating the ban on material support but who can be subject to warrantless surveillance, indefinite detention, and summary execution by drone. But don't worry: The Obama administration is providing all the process it believes is due.


01 Oct 21:57

Want to be unhappy? Trying to be happy will do it!

by mdbownds@wisc.edu (Deric Bownds)
I'm finding the "Anxiety" topic in the Opinionator series at the NYTimes to be a real treat. This entry by British expatriate Ruth Whippman brings back memories of a my signing on several years ago to be a talking head neuroscience expert on the  California "Make Me Happy!" Radio Show (I don't think they were all that pleased with their dyspeptic guest!). Whippman notes the American obsession with, and anxiety over, being "happy", and contrasts this with the attitudes of more stoic Britishers:
Happiness in America has become the overachiever's ultimate trophy. A vicious trump card, it outranks professional achievement and social success, family, friendship and even love…this elusive MacGuffin is creating a nation of nervous wrecks. Despite being the richest nation on earth, the United States is, according to the World Health Organization, by a wide margin, also the most anxious, with nearly a third of Americans likely to suffer from an anxiety problem in their lifetime. America's precocious levels of anxiety are not just happening in spite of the great national happiness rat race, but also perhaps, because of it.
The British are generally uncomfortable around the subject, and as a rule, don't subscribe to the happy-ever-after. It's not that we don't want to be happy, it just seems somehow embarrassing to discuss it, and demeaning to chase it, like calling someone moments after a first date to ask them if they like you….Even the recent grand spectacle of the London 2012 Olympic Games told this tale. The opening ceremony, traditionally a sparklefest of perkiness, was, with its suffragist and trade unionists, mainly a celebration of dissent, or put less grandly, complaint…Our queen, despite the repeated presence of a stadium full of her subjects urging in song that she be both happy and glorious, could barely muster a smile, staring grimly through her eyeglasses and clutching her purse on her lap as if she might be mugged.
Cynicism is the British shtick. When happiness does come our way, it is entirely without effort, as unmeritocratic as a hereditary peerage. By contrast, in America, happiness is work. Intense, nail-biting work, slogged out in motivational seminars and therapy sessions, meditation retreats and airport bookstores. For the left there's yoga, for the right, there's Jesus. For no one is there respite…The people taking part in "happiness pursuits," as a rule, don't seem very happy…The happy person would be more likely to be off doing something fun, like sitting in the park drinking.
Happiness should be serendipitous, a by-product of a life well lived, and pursuing it in a vacuum doesn't really work. This is borne out by a series of slightly depressing statistics. The most likely customer of a self-help book is a person who has bought another self-help book in the last 18 months. The General Social Survey, a prominent data-based barometer of American society, shows little change in happiness levels since 1972, when such records began. Every year, with remarkable consistency, around 33 percent of Americans report that they are "very happy." It's a fair chunk, but a figure that remains surprisingly constant, untouched by the uptick in Eastern meditation or evangelical Christianity, by Tony Robbins or Gretchen Rubin or attachment parenting. For all the effort Americans are putting into happiness, they are not getting any happier. It is not surprising, then, that the search itself has become a source of anxiety.
So here's a bumper sticker: despite the glorious weather and spectacular landscape, the people of California are probably less happy and more anxious than the people of Grimsby. So they may as well stop trying so hard.
24 Sep 22:58

Strange-face illusions during inter-subjective gazing.

Conscious Cogn. 2012 Sep 12; Caputo GBIn normal observers, gazing at one's own face in the mirror for a few minutes, at a low illumination level, triggers the perception of strange faces, a new visual illusion that has been named 'strange-face in the mirror'. Individuals see huge distortions of their own faces, but they often see monstrous beings, archetypal faces, faces of relatives and deceased, and animals. In the experiment described here, strange-face illusions were perceived when two individuals, in a dimly lit room, gazed at each other in the face. Inter-subjective gazing compared to mirror-gazing produced a higher number of different strange-faces. Inter-subjective strange-face illusions were always dissociative of the subject's self and supported moderate feeling of their reality, indicating a temporary lost of self-agency. Unconscious synchronization of event-related responses to illusions was found between members in some pairs. Synchrony of illusions may indicate that unconscious response-coordination is caused by the illusion-conjunction of crossed dissociative strange-faces, which are perceived as projections into each other's visual face of reciprocal embodied representations within the pair. Inter-subjective strange-face illusions may be explained by the subject's embodied representations (somaesthetic, kinaesthetic and motor facial pattern) and the other's visual face binding. Unconscious facial mimicry may promote inter-subjective illusion-conjunction, then unconscious joint-action and response-coordination.
23 Sep 14:51

A football game has only 11 minutes of action

by Minnesotastan
From a 2010 article in the WSJ (the numbers might have changed a bit since then):
According to a Wall Street Journal study of four recent broadcasts, and similar estimates by researchers, the average amount of time the ball is in play on the field during an NFL game is about 11 minutes...

So what do the networks do with the other 174 minutes in a typical broadcast? Not surprisingly, commercials take up about an hour. As many as 75 minutes, or about 60% of the total air time, excluding commercials, is spent on shots of players huddling, standing at the line of scrimmage or just generally milling about between snaps. In the four broadcasts The Journal studied, injured players got six more seconds of camera time than celebrating players. While the network announcers showed up on screen for just 30 seconds, shots of the head coaches and referees took up about 7% of the average show...
This is why the only way I watch football nowadays is by using a DVR and speeding through the game (and past the commercials).

Reposted from 2012 to add news of developments in 2017:
"It has been an effort for a long period of time. We've talked about the length of the game," [NFL Commissioner] Goodell said. "This effort's not as focused on the length of the game. This is focused on what's happening outside the plays -- how fast we get the ball set, the number of breaks, the number of intrusions -- so that fans can focus on the action."

With all this talk about making the game faster for fans, what would Goodell consider the ideal length of a broadcast?
"We (were at) 3:07 and change (last season), down about a minute," Goodell said. "We think we could probably get pretty close to five minutes of downtime out of the game, so that would bring you somewhere in the 3:02 range. That would be very successful if we could get to that point. But, again, not just the length. We want to make sure we are taking the right things out of the game -- the things that are not compelling to our fans."
Clueless.  The idea that cutting 5 minutes out of a 3-hour broadcast will satisfy fans' frustrations shows that viewer interests don't even begin to compete with advertiser's interests.
12 Sep 04:08

Thomas Szasz: How and Why the Great Libertarian Psychiatrist Thought What He Did

by Brian Doherty

Jacob Sullum and Jesse Walker have both done great jobs summing up the importance of Szasz; I have always found his own thoughts and expressions the best way to understand him. He was, in my judgement, one of the smartest and most thorough defenders of autonomy and liberty of our time, fighting against both his profession, most of the world, and often his own fellow libertarians, and succeding at a higher level than most (Szasz was actually a public intellectual of mass popularity in the 1960s/early 1970s.)

Herewith, a sampling from some of my own previous writings about Szasz, mostly quoting him.

From a 1999 review essay for Feed magazine:

Szasz says that most so-called mental illnesses are not what the psychiatric profession maintains, and that fact is of great socio-political and ethical importance....

Szasz says the category of "mental illness" turns willed behavior into a disease, taking away both rights and responsibilities from the actor just because his actions strikes a doctor, family member, or judge as inexplicably bizarre and strange. In pragmatic terms, Szasz avers, "incarcerating innocent persons in mental hospitals and freeing guilty persons from prison... continue to be the psychiatrist's two most important social functions." He takes a cui bono? approach, asking what the psychiatric profession gains from the idea of mental illness (prestige, power, money) and what the patient gains (exculpation for bad actions or crime, relief from responsibility). Szasz is politically appalled by the coercion inherent in the modern psychiatric enterprise, and always credits even the most seemingly mad with humanity and intentionality. On the contrary, psychiatrists rarely credit the lunatic with having any sense or rationality behind his actions -- even when it's clear that there is some rational goal in mind. "A berserk lunatic may claim to be Jesus or kill his wife," Szasz writes. "The point of such a person's behavior, I dare say, is to be revered like Jesus or be rid of his wife. (Why a person chooses such ends and means is another question, the answer to which is often easily obtained by asking him.)"

From my 2007 book Radicals for Capitalism:

The innovation—or semantic trick, as Szasz would have it—of classical psychoanalysis was turning faking an illness into an illness in and of itself. The human capacity for deception is central to Szasz’s intellectual program. Human beings lie; and many a so-called insane delusion, such as voices in the head advising one to commit heinous acts, are, Szasz maintained, best understood as lies—often strategic ones....

Szasz...recalls that “I was not about to tell him that the persons he called ‘seriously ill patients’ I regarded as persons deprived of liberty by psychiatrists.”...He later wrote that “psychiatric training is, above all else, a ritualized indoctrination into the theory and practice of psychiatric violence. The disastrous effects of this process on the patients are obvious enough; though less evident, its consequences for the physician are often equally tragic.”

....To Szasz, psychoanalysis proper had nothing to do with medicine. It was conversation, with one person paying the other. “The psychiatrist has only one duty: to keep his mouth shut outside the room and maintain total confidentiality. It has nothing to do with disease. It has to do with human problems."...

Szasz fought for the specific liberty of specific patients:

“When I began to publish on the civil rights of mental patients, some of this hit the papers, The New York Times. I began to get invitations from patients and lawyers—‘I have this client locked up for 10 years and he hasn’t done anything. He’s been in long enough. Can you get him out?’” Szasz began testifying on behalf of imprisoned mental patients—some alternately hilarious and harrowing transcripts from those court cases are in his book Psychiatric Justice (1965)—though he rarely succeeded in winning anyone’s freedom. He’d find himself, he recalled, “in the courtroom in front of some very nice judge who said something like, ‘Szasz, how can you say that he should be out when six of his doctors say he should be in?’ I said, ‘Your honor, those are not his doctors. Those are his adversaries. He wants his freedom. I am the one that he calls his doctor.’”....

Szasz thought he saw the underlying truth of psychiatry others missed, or wanted to miss:

[Szasz] considers his heterodox positions pure common sense, a common sense marred by the power-grabbing pretensions of psychiatrists and the government-psychiatric establishment. Those pretensions have been embraced by a credulous populace all too ready to believe that people should be relieved of both responsibility and liberty whenever it became convenient for either the state or any relative or caretaker troubled by the so-called mentally ill. From the very beginning, Szasz recognized that psychiatry wasn’t really about what it purported to be about.

“What is the thing itself that psychiatrists describe, debate, diagnose, and treat?” Szasz asked. “The psychiatrist says it is mental illness, which, he now quickly adds, is the name of neurochemical lesions of the brain. I say it is conflict and coercion and the rules that regulate the psychiatrist’s power and privileges and the patient’s rights and responsibilities. The former perspective leads to an analysis of psychiatry in terms of illness and treatment, medical theory and therapeutic practice, while [my] perspective leads to an analysis in terms of coercion and contract, the exercise of power and the efforts to limit it, in short, political theory and legal practice.” He believed that psychiatry was more properly conceived as an ethical and political field—the arena of human troubles, communication, and conflict—than as a medical science. Psychiatry was rife with “hidden agendas of domination and submission concealed by a rhetoric of disease and treatment.”....

Szasz was anti-coercive-psychiatry, but not a cliched "anti-psychiatrist."

While defending the rights of mental patients not to be treated or imprisoned against their will, Szasz was dismissive of the “anti-psychiatry” movement and its figurehead R.D. Laing, with whom Szasz was often mistakenly conflated in the late ‘60s and early ‘70s. Szasz had little sympathy with the Laingian view that saw the so-called insane as in fact victims of an insane society—or going through an understandable reaction to that insane society—or visionaries taking a valuable “journey through madness.”

“I insist,” Szasz wrote, “that schizophrenia is no more a journey through madness than it is a disease of the brain. Both of these statements assert literalized metaphors. Of course schizophrenia may be said to be like a journey or like a disease; but it is also like many other conditions or situations; for example, being childish, aimless, useless, and homeless, or being angry, obstreperous, conceited, or selfish.”

Laingian assessments of the so-called insane, then, were in most cases higher than Szasz’s, who is above all a moralist, and not a groovy admirer of alternative lifestyles. (A student skit at his university joked about the “Szasz Diagnostic Manual” which had two categories: “crook” and “bum.”) The antipsychiatrists, to Szasz, were just as paternalistic and anti-individualist as their opponents, merely in the opposite direction...

Szasz is perplexed that any part of the psychiatric industry sees him as anything other than a bitter enemy. He speculates that those of his colleagues who accept him as a friendly and welcome addition to the scholarly debate “just don’t give serious enough thought to this to either agree or dismiss it and dismiss me as completely wrong. They just write me off as ‘interesting.’"

Szasz had unique things to say to libertarians:

He has analogized a sane human life to a statue carved out of marble. Although we may all metaphorically have a chunk of marble at birth, we don’t all automatically have the statue, as standard mental health professionals seem to think we ought; nor does the lack of a statue mean a repressive culture has smashed ours. It means we haven’t done the work to sculpt it. Szasz is the libertarian movement’s most stoic exponent, hoping for a fully free and responsible culture but painfully mindful that it may be impossible—for reasons that don’t necessarily have to do with the outward tyranny of the state.

Szasz was also very personally gracious to this young reporter and fan, giving me time and attention above the call of duty when we interacted professionally. He was a model public intellectual and a decent and brave man.