Shared posts

28 Jun 08:00

Simultaneous Measurement of Complementary Observables with Compressive Sensing

by Gregory A. Howland, James Schneeloch, Daniel J. Lum, and John C. Howell

Author(s): Gregory A. Howland, James Schneeloch, Daniel J. Lum, and John C. Howell

Coincident high-resolution position and momentum imaging of optical photons is achieved using a sequence of weak and strong measurements.

[Phys. Rev. Lett. 112, 253602] Published Thu Jun 26, 2014

23 Jun 05:24

When being a control-freak doesn't help....

by mdbownds@wisc.edu (Deric Bownds)
Bocanegra and Hommel note limits to the usefulness of cognitive control, showing, in particular, how overcontrol (induced by task instructions) can prevent the otherwise automatic exploitation of statistical stimulus characteristics needed to optimize behavior. They describe how they set up the experiment:
Participants performed a two-alternative forced-choice task on a foveally presented stimulus that could vary on a subset of binary perceptual features, such as color (red, green), shape (diamond, square), size (large, small), topology (open, closed), and location (up, down). Unbeknownst to the participants, we manipulated the statistical informativeness of an additional feature that was not part of the task, such that this feature always predicted the correct response in one condition (the predictive condition) but not in the other condition (the baseline condition). Because the cognitive system is known to exploit statistical stimulus-response contingencies automatically, performance was expected to be better in the predictive than in the baseline condition.
We embedded these predictive and baseline conditions into two different tasks, which we thought would induce different cognitive-control states. The control task included instructions intended to emphasize the need for top-down control: Participants were instructed to classify the stimulus according to a feature-conjunction rule (e.g., size and topology: left response key for large and open or small and closed shapes, right response key for small and open or large and closed shapes). The automatic task included instructions intended to deemphasize the need for control: Participants were instructed to classify the stimulus according to a single feature (e.g., shape: left response key for a diamond and right response key for a square). In the automatic task, the features were mapped consistently on responses and thus allowed automatic visuomotor translation. In contrast, the stimulus-response mapping in the control task required the attention-demanding integration of two features before the response could be determined.
As expected, the predictive feature improved performance when participants performed the task automatically. Counterintuitively, however, the predictive feature impaired performance when subjects were performing the exact same task in a top-down, controlled manner.
Their abstract:
In order to engage in goal-directed behavior, cognitive agents have to control the processing of task-relevant features in their environments. Although cognitive control is critical for performance in unpredictable task environments, it is currently unknown how it affects performance in highly structured and predictable environments. In the present study, we showed that, counterintuitively, top-down control can impair and interfere with the otherwise automatic integration of statistical information in a predictable task environment, and it can render behavior less efficient than it would have been without the attempt to control the flow of information. In other words, less can sometimes be more (in terms of cognitive control), especially if the environment provides sufficient information for the cognitive system to behave on autopilot based on automatic processes alone.
22 Jun 02:01

Takens' Embedding and Riemannian preconditioning

by Igor
If wish I had known about the Johnson-Lindenstrauss Lemma or the Taken's embedding theorem or even surfing on manifolds to do things faster much earlier than I did. For the sake of putting things in context, here is the connection between compressive sensing and the Johnson-Lindenstrauss lemma (from The Johnson-Lindenstrauss Lemma Meets Compressed Sensing):

...In Compressed Sensing (CS) [2, 7], for example, a random projection of a highdimensional but sparse signal vector onto a lower-dimensional space has been shown, with high probability, to contain enough information to enable signal reconstruction with small or zero error. Random projections also play a fundamental role in the study of point clouds in high-dimensional spaces. The Johnson-Lindenstrauss (JL) lemma [12], for example, shows that with high probability the geometry of a point cloud is not disturbed by certain Lipschitz mappings onto a space of dimension logarithmic in the number of points. The statement and proofs of the JL-lemma have been simplified considerably by using random linear projections and concentration inequalities [1]....

Equally, why should we care about Takens' embedding theorem ? here is some insight (from the first paper) as to why we should care:

"...The underlying problem is that while Takens’ theorem guarantees the preservation of the attractor’s topology, it does not guarantee that the geometry of the attractor is also preserved. To be precise, Takens’ result guarantees that two points on the attractor do not map to the same point in the reconstruction space, but there are no guarantees that close points on the attractor remain close under this mapping (or far points remain far). Consequently, relatively small imperfections could have arbitrarily large, unwanted effects when the delay coordinate map is used in applications..."

The second paper takes on a somewhat related route, where, beyond defining specific metrics for specific objects like Toeplitz or psd matrices and whatnot, the next step is to use those to perform some computations faster ( see also some of these recent comments)



Takens' Embedding Theorem asserts that when the states of a hidden dynamical system are confined to a low-dimensional attractor, complete information about the states can be preserved in the observed time-series output through the delay coordinate map. However, the conditions for the theorem to hold ignore the effects of noise and time-series analysis in practice requires a careful empirical determination of the sampling time and number of delays resulting in a number of delay coordinates larger than the minimum prescribed by Takens' theorem. In this paper, we use tools and ideas in Compressed Sensing to provide a first theoretical justification for the choice of the number of delays in noisy conditions. In particular, we show that under certain conditions on the dynamical system, measurement function, number of delays and sampling time, the delay-coordinate map can be a stable embedding of the dynamical system's attractor.


Riemannian preconditioning by Bamdev Mishra, Rodolphe Sepulchre
The paper exploits a basic connection between sequential quadratic programming and Riemannian gradient optimization to address the general question of selecting a metric in Riemannian optimization. The proposed method is shown to be particularly insightful and efficient in quadratic optimization with orthogonality and/or rank constraints, which covers most current applications of Riemannian optimization in matrix manifolds.

12 Jun 16:53

Categorification, step 1

by Peter Cameron

Today at the St Petersburg meeting, Igor Frenkel talked about categorification. He explained that there are five levels (maybe more!) and one has to take certain steps between them; he illustrated with an example, where level 0 was Jacobi’s Triple Product Identity and level 5 was four-dimensional quantum Yang–Mills.

He did say,

Forget about categorification if you don’t want to hear this word,

and it is certainly not an attractive word.

But I, like a stubborn mule, found myself unable to cross the Pons Asinorum from Level 0 to Level 1, since the landscape along the way was too attractive. I have some questions about this, which I will describe briefly here. I may try to write this out at greater length sometime.

Numbers and vector spaces

Step 1, in essence, consists in replacing numbers (non-negative integers) with vector spaces over a field, which is usually taken to be the field of complex numbers (but undoubtedly different things will happen over other fields). The number n is replaced by an n-dimensional vector space over C.

Addition and multiplication of numbers is replaced by direct sum and tensor product of vector space; the dimensions do indeed behave correctly.

An equation like ab+c = 0 is replaced by a short exact sequence

{0} → A → B → C → {0};

this means that the image of each map is equal to the kernel of the next. Again the dimensions behave correctly. But we see already that the process is not purely mechanical, since the reverse of the short exact sequence (which could be realised by maps of the dual spaces) may not be the thing that naturally occurs.

For example, Euler’s polyhedral formula (written in the form 1−V+EF+1 = 0) can be realised in this way: first orient the edges and faces of the polyhedron; now replace V,E,F by vector spaces of functions from vertices, edges and faces to C; and then use “coboundary maps” to make the sequence. (For example, from V to E, we replace a function f on vertices by a function df on edges, where df(v,w) = f(w)−f(v).) Now the proof of Euler’s formula becomes the proof of exactness of this sequence.

Formal power series

A formal power series ∑anxn has to be treated a bit differently. We cannot substitute a natural number for x, since all but the tamest such sequences have radius of convergence smaller than 1. But, as in combinatorial enumeration, we regard the powers of x as markers.

Thus, the interpretation of the series is a graded vector space ⊕An, where An is a vector space of dimension an.

Now direct sum or tensor product of graded vector spaces (with the usual conventions about grading, that is, the product of elements of degrees k and l has degree k+l), correspond to the sum or product of the formal power series.

Problem What if the graded vector space is actually a graded algebra?

In particular, the formal power series 1 and x corresponds to 1-dimensional vector spaces with degree 0 and 1 respectively.

Invariants

Let G be a finite permutation group on the set {1,…,n}. Then G has a natural action on a vector space V of dimension n, by permuting the vectors of a basis. An invariant of G is simply a vector fixed by G; such a vector must have coordinates which are constant on each G-orbit, and so the space of invariants of G (written VG) has dimension equal to the number of orbits of G. So we have taken the first step towards categorifying orbit-counting.

More generally, let W be a graded vector space. Then there is a natural action of G on the direct sum V = Wn of n copies of G. How do we “count” invariants here?

The answer lies in Pólya theory. Associated with G is a multivariate polynomial called the cycle index of G. This is the polynomial in indeterminates s1,…sn constructed as follows: for each element g of the group, having ci cycles of length i for each i, we form a monomial by raising si to the power ci, multiplying all these together; then sum these monomials over all group elements, and divide by the order of the group. This is denoted by Z(G).

Suppose that we have a collection of “figures”, each with a non-negative integer “weight”, so that there are only finitely many (say ai) figures of weight i for each i. The figures can be represented by a figure-counting series A(x) = ∑aixi. Now a “function” is a map from the set {1,…n} to the set of figures (essentially it puts a figure at every point of this set). G permutes the functions by moving their arguments. We are interested in counting the orbits of G on functions of given total weight (the sum of the weights of their values). If bi is the number of orbits, then the function-counting series is B(x) = ∑bixi.

Now Pólya’s Theorem asserts that B(x) is obtained from Z(G) by substituting A(xi) for si, for each i.

Now (I think), if we replace the figure-counting series by a graded vector space W, and let G act on the direct sum V of n copies of V, then the function-counting series corresponds to the space of invariants of G in V.

Of course, permutation actions are a special case of linear actions.

Problem. What happens for linear actions? Do we replace Pólya’s theorem by Molien’s?

Oligomorphic groups

A permutation group G on an infinite set is called oligomorphic if it has only finitely many (say an) orbits on the set of all n-element subsets of the permutation domain.

The definition of cycle index breaks down completely for oligomorphic permutation groups. However, one can define a modified cycle index for an oligomorphic group; this is obtained by taking orbit representatives for the orbits of G on the finite subsets of its domain, for each orbit representative take the cycle index of the finite group induced on this set by its setwise stabiliser in G, and then summing. It is easy to see that we obtain a formal power series in the infinitely many indeterminates s1,s2,….

A finite permutation group is a special case of an oligomorphic group; according to the Shift Theorem, its modified cycle index can be obtained from its ordinary cycle index by replacing si by si+1, for each i.

Many substitution results hold. For example, the power series ∑anxn is obtained from the modified cycle index by substituting xi for si, for each i. There are also rules for direct and wreath products.

There is at least the possibility of turning all this into graded vector spaces. Let Vi be the vector space of all functions from the set of G-orbits on i-sets to C, and V the graded vector space having these spaces as their homogeneous components. Then the vector space V “represents” the orbit-counting power series.

The vector space V has a natural algebra structure, which I won’t describe here. In fact, though results about the rate of growth of the numbers an tend to be proved by combinatorial methods (chiefly by Dugald Macpherson), I have had a small amount of success in proving smoothness results using algebraic methods, based on the fact that the multiplication in the algebra gives a map from VmVn to Vm+n.

Problem What are the vector space operations corresponding to direct and wreath product of oligomorphic groups?

Problem Is there a linear group analogue of oligomorphic permutation groups, for which a similar theory can be developed?

I think that is enough problems to be going on with!


11 Jun 22:16

Alexander Shulgin, RIP

by Jacob Sullum

Alexander Shulgin, who died this week at the age of 88, was a remarkable man who combined an intense curiosity about altered states of consciousness with amazing chemical creativity and scientific rigor. Over the years Shulgin synthesized hundreds of psychoactive compounds that he carefully tested on himself, his wife, Ann, and a small circle of friends—a process he described in his 1991 book PIKHAL: A Chemical Love Story, a 978-page tome that includes notes on the production and effects of 179 such chemicals along with a personal and professional memoir. (The title stands for "Phenethylamines I Have Known and Loved"; the sequel was called TIKHAL, for "Tryptamines I Have Known and Loved.") Perhaps best known as a popularizer (though not the creator) of MDMA, which he said "enabled me to see out, and to see my own insides, without reservations," Shulgin embodied an open-minded yet responsible approach to drugs that should be a model for psychonauts as well as the politicians who vainly try to control them.

"Every drug, legal or illegal, provides some reward," Shulgin wrote in the introduction to PIKHAL. "Every drug presents some risk. And every drug can be abused. Ultimately, in my opinion, it is up to each of us to measure the reward against the risk and decide which outweighs the other." His great passion was for psychedelics, the "mind-manifesting" drugs with effects similar to those of mescaline, psilocybin, and LSD, which he saw as "treasures" that "can provide access to the parts of us that have answers," facilitating "exploration of this interior world" and "insights into its nature." It amazed him that legislators and regulators would presume to intrude into this deeply personal realm, especially in a society that claims to respect privacy, freedom of inquiry, and freedom of conscience.

"Our generation is the first, ever, to have made the search for self-awareness a crime, if it is done with the use of plants or chemical compounds as the means of opening the psychic doors," Shulgin wrote. "How is it...that the leaders of our society have seen fit to try to eliminate this one very important means of learning and self-discovery, this means which has been used, respected, and honored for thousands of years, in every human culture of which we have a record?"

That remains a bit of a puzzle, even to people who have studied the series of moral panics that comprise the history of American drug policy. But by highlighting the profound, life-enhancing potential of forbidden intoxicants without denying their hazards, Shulgin boldly pointed the way to a more tolerant alternative.

11 Jun 21:37

Anomalous Transfer of Syntax between Languages

by Vaughan-Evans, A., Kuipers, J. R., Thierry, G., Jones, M. W.

Each human language possesses a set of distinctive syntactic rules. Here, we show that balanced Welsh-English bilinguals reading in English unconsciously apply a morphosyntactic rule that only exists in Welsh. The Welsh soft mutation rule determines whether the initial consonant of a noun changes based on the grammatical context (e.g., the feminine noun cath—"cat" mutates into gath in the phrase y gath—"the cat"). Using event-related brain potentials, we establish that English nouns artificially mutated according to the Welsh mutation rule (e.g., "goncert" instead of "concert") require significantly less processing effort than the same nouns implicitly violating Welsh syntax. Crucially, this effect is found whether or not the mutation affects the same initial consonant in English and Welsh, showing that Welsh syntax is applied to English regardless of phonological overlap between the two languages. Overall, these results demonstrate for the first time that abstract syntactic rules transfer anomalously from one language to the other, even when such rules exist only in one language.

31 May 16:14

Post-linear Schwarzschild solution in harmonic coordinates: Elimination of structure-dependent terms

by Sergei A. Klioner and Michael Soffel
Nosimpler

Shared for "post-linear", which is a word we should adapt for our own purposes.

Author(s): Sergei A. Klioner and Michael Soffel

This paper deals with a special kind of problems that appear in solutions of Einstein’s field equations for extended bodies: many structure-dependent terms appear in intermediate calculations that cancel exactly in virtue of the local equations of motion or can be eliminated by appropriate gauge tra...

[Phys. Rev. D 89, 104056] Published Tue May 27, 2014

31 May 16:14

Is there signal in the noise?

by Alexander S Ecker
Nosimpler

Take that, noise-worshippers.

Nature Neuroscience 17, 750 (2014). doi:10.1038/nn.3722

Authors: Alexander S Ecker & Andreas S Tolias

A study now shows that variability in neuronal responses in the visual system mainly arises from slow fluctuations in excitability, presumably caused by factors of nonsensory origin, such as arousal, attention or anesthesia.

30 May 17:50

Universal features in the energetics of symmetry breaking

by É. Roldán

Nature Physics 10, 457 (2014). doi:10.1038/nphys2940

Authors: É. Roldán, I. A. Martínez, J. M. R. Parrondo & D. Petrov

28 May 18:48

Future Hollywood movies superstar

by Minnesotastan

John Wayne.

Via Reddit.
26 May 16:57

Nurses complain about algorithms

by Tyler Cowen
Nosimpler

Mo automation, mo problems.

In response [to the rise of diagnostic algorithms], NNU [National Nurses United] has launched a major campaign featuring radio ads from coast to coast, video, social media, legislation, rallies, and a call to the public to act, with a simple theme – “when it matters most, insist on a registered nurse.”  The ads were created by North Woods Advertising and produced by Fortaleza Films/Los Angeles.  Additional background can be found at http://www.insistonanrn.org.

Here is the link.  Here is an MP3 of the ad.  Remarkable, do give it a listen.  It has numerous excellent lines such as “Algorithms are simple mathematical formulas that nobody understands.”

For the pointer I thank Eric Jonas.

24 May 18:17

Why I Am Not An Integrated Information Theorist (or, The Unconscious Expander)

by Scott

Happy birthday to me!

Recently, lots of people have been asking me what I think about IIT—no, not the Indian Institutes of Technology, but Integrated Information Theory, a widely-discussed “mathematical theory of consciousness” developed over the past decade by the neuroscientist Giulio Tononi.  One of the askers was Max Tegmark, who’s enthusiastically adopted IIT as a plank in his radical mathematizing platform (see his paper “Consciousness as a State of Matter”).  When, in the comment thread about Max’s Mathematical Universe Hypothesis, I expressed doubts about IIT, Max challenged me to back up my doubts with a quantitative calculation.

So, this is the post that I promised to Max and all the others, about why I don’t believe IIT.  And yes, it will contain that quantitative calculation.

But first, what is IIT?  The central ideas of IIT, as I understand them, are:

(1) to propose a quantitative measure, called Φ, of the amount of “integrated information” in a physical system (i.e. information that can’t be localized in the system’s individual parts), and then

(2) to hypothesize that a physical system is “conscious” if and only if it has a large value of Φ—and indeed, that a system is more conscious the larger its Φ value.

I’ll return later to the precise definition of Φ—but basically, it’s obtained by minimizing, over all subdivisions of your physical system into two parts A and B, some measure of the mutual information between A’s outputs and B’s inputs and vice versa.  Now, one immediate consequence of any definition like this is that all sorts of simple physical systems (a thermostat, a photodiode, etc.) will turn out to have small but nonzero Φ values.  To his credit, Tononi cheerfully accepts the panpsychist implication: yes, he says, it really does mean that thermostats and photodiodes have small but nonzero levels of consciousness.  On the other hand, for the theory to work, it had better be the case that Φ is small for “intuitively unconscious” systems, and only large for “intuitively conscious” systems.  As I’ll explain later, this strikes me as a crucial point on which IIT fails.

The literature on IIT is too big to do it justice in a blog post.  Strikingly, in addition to the “primary” literature, there’s now even a “secondary” literature, which treats IIT as a sort of established base on which to build further speculations about consciousness.  Besides the Tegmark paper linked to above, see for example this paper by Maguire et al., and associated popular article.  (Ironically, Maguire et al. use IIT to argue for the Penrose-like view that consciousness might have uncomputable aspects—a use diametrically opposed to Tegmark’s.)

Anyway, if you want to read a popular article about IIT, there are loads of them: see here for the New York Times’s, here for Scientific American‘s, here for IEEE Spectrum‘s, and here for the New Yorker‘s.  Unfortunately, none of those articles will tell you the meat (i.e., the definition of integrated information); for that you need technical papers, like this or this by Tononi, or this by Seth et al.  IIT is also described in Christof Koch’s memoir Consciousness: Confessions of a Romantic Reductionist, which I read and enjoyed; as well as Tononi’s Phi: A Voyage from the Brain to the Soul, which I haven’t yet read.  (Koch, one of the world’s best-known thinkers and writers about consciousness, has also become an evangelist for IIT.)

So, I want to explain why I don’t think IIT solves even the problem that it “plausibly could have” solved.  But before I can do that, I need to do some philosophical ground-clearing.  Broadly speaking, what is it that a “mathematical theory of consciousness” is supposed to do?  What questions should it answer, and how should we judge whether it’s succeeded?

The most obvious thing a consciousness theory could do is to explain why consciousness exists: that is, to solve what David Chalmers calls the “Hard Problem,” by telling us how a clump of neurons is able to give rise to the taste of strawberries, the redness of red … you know, all that ineffable first-persony stuff.  Alas, there’s a strong argument—one that I, personally, find completely convincing—why that’s too much to ask of any scientific theory.  Namely, no matter what the third-person facts were, one could always imagine a universe consistent with those facts in which no one “really” experienced anything.  So for example, if someone claims that integrated information “explains” why consciousness exists—nope, sorry!  I’ve just conjured into my imagination beings whose Φ-values are a thousand, nay a trillion times larger than humans’, yet who are also philosophical zombies: entities that there’s nothing that it’s like to be.  Granted, maybe such zombies can’t exist in the actual world: maybe, if you tried to create one, God would notice its large Φ-value and generously bequeath it a soul.  But if so, then that’s a further fact about our world, a fact that manifestly couldn’t be deduced from the properties of Φ alone.  Notice that the details of Φ are completely irrelevant to the argument.

Faced with this point, many scientifically-minded people start yelling and throwing things.  They say that “zombies” and so forth are empty metaphysics, and that our only hope of learning about consciousness is to engage with actual facts about the brain.  And that’s a perfectly reasonable position!  As far as I’m concerned, you absolutely have the option of dismissing Chalmers’ Hard Problem as a navel-gazing distraction from the real work of neuroscience.  The one thing you can’t do is have it both ways: that is, you can’t say both that the Hard Problem is meaningless, and that progress in neuroscience will soon solve the problem if it hasn’t already.  You can’t maintain simultaneously that

(a) once you account for someone’s observed behavior and the details of their brain organization, there’s nothing further about consciousness to be explained, and

(b) remarkably, the XYZ theory of consciousness can explain the “nothing further” (e.g., by reducing it to integrated information processing), or might be on the verge of doing so.

As obvious as this sounds, it seems to me that large swaths of consciousness-theorizing can just be summarily rejected for trying to have their brain and eat it in precisely the above way.

Fortunately, I think IIT survives the above observations.  For we can easily interpret IIT as trying to do something more “modest” than solve the Hard Problem, although still staggeringly audacious.  Namely, we can say that IIT “merely” aims to tell us which physical systems are associated with consciousness and which aren’t, purely in terms of the systems’ physical organization.  The test of such a theory is whether it can produce results agreeing with “commonsense intuition”: for example, whether it can affirm, from first principles, that (most) humans are conscious; that dogs and horses are also conscious but less so; that rocks, livers, bacteria colonies, and existing digital computers are not conscious (or are hardly conscious); and that a room full of people has no “mega-consciousness” over and above the consciousnesses of the individuals.

The reason it’s so important that the theory uphold “common sense” on these test cases is that, given the experimental inaccessibility of consciousness, this is basically the only test available to us.  If the theory gets the test cases “wrong” (i.e., gives results diverging from common sense), it’s not clear that there’s anything else for the theory to get “right.”  Of course, supposing we had a theory that got the test cases right, we could then have a field day with the less-obvious cases, programming our computers to tell us exactly how much consciousness is present in octopi, fetuses, brain-damaged patients, and hypothetical AI bots.

In my opinion, how to construct a theory that tells us which physical systems are conscious and which aren’t—giving answers that agree with “common sense” whenever the latter renders a verdict—is one of the deepest, most fascinating problems in all of science.  Since I don’t know a standard name for the problem, I hereby call it the Pretty-Hard Problem of Consciousness.  Unlike with the Hard Hard Problem, I don’t know of any philosophical reason why the Pretty-Hard Problem should be inherently unsolvable; but on the other hand, humans seem nowhere close to solving it (if we had solved it, then we could reduce the abortion, animal rights, and strong AI debates to “gentlemen, let us calculate!”).

Now, I regard IIT as a serious, honorable attempt to grapple with the Pretty-Hard Problem of Consciousness: something concrete enough to move the discussion forward.  But I also regard IIT as a failed attempt on the problem.  And I wish people would recognize its failure, learn from it, and move on.

In my view, IIT fails to solve the Pretty-Hard Problem because it unavoidably predicts vast amounts of consciousness in physical systems that no sane person would regard as particularly “conscious” at all: indeed, systems that do nothing but apply a low-density parity-check code, or other simple transformations of their input data.  Moreover, IIT predicts not merely that these systems are “slightly” conscious (which would be fine), but that they can be unboundedly more conscious than humans are.

To justify that claim, I first need to define Φ.  Strikingly, despite the large literature about Φ, I had a hard time finding a clear mathematical definition of it—one that not only listed formulas but fully defined the structures that the formulas were talking about.  Complicating matters further, there are several competing definitions of Φ in the literature, including ΦDM (discrete memoryless), ΦE (empirical), and ΦAR (autoregressive), which apply in different contexts (e.g., some take time evolution into account and others don’t).  Nevertheless, I think I can define Φ in a way that will make sense to theoretical computer scientists.  And crucially, the broad point I want to make about Φ won’t depend much on the details of its formalization anyway.

We consider a discrete system in a state x=(x1,…,xn)∈Sn, where S is a finite alphabet (the simplest case is S={0,1}).  We imagine that the system evolves via an “updating function” f:Sn→Sn. Then the question that interests us is whether the xi‘s can be partitioned into two sets A and B, of roughly comparable size, such that the updates to the variables in A don’t depend very much on the variables in B and vice versa.  If such a partition exists, then we say that the computation of f does not involve “global integration of information,” which on Tononi’s theory is a defining aspect of consciousness.

More formally, given a partition (A,B) of {1,…,n}, let us write an input y=(y1,…,yn)∈Sn to f in the form (yA,yB), where yA consists of the y variables in A and yB consists of the y variables in B.  Then we can think of f as mapping an input pair (yA,yB) to an output pair (zA,zB).  Now, we define the “effective information” EI(A→B) as H(zB | A random, yB=xB).  Or in words, EI(A→B) is the Shannon entropy of the output variables in B, if the input variables in A are drawn uniformly at random, while the input variables in B are fixed to their values in x.  It’s a measure of the dependence of B on A in the computation of f(x).  Similarly, we define

EI(B→A) := H(zA | B random, yA=xA).

We then consider the sum

Φ(A,B) := EI(A→B) + EI(B→A).

Intuitively, we’d like the integrated information Φ=Φ(f,x) be the minimum of Φ(A,B), over all 2n-2 possible partitions of {1,…,n} into nonempty sets A and B.  The idea is that Φ should be large, if and only if it’s not possible to partition the variables into two sets A and B, in such a way that not much information flows from A to B or vice versa when f(x) is computed.

However, no sooner do we propose this than we notice a technical problem.  What if A is much larger than B, or vice versa?  As an extreme case, what if A={1,…,n-1} and B={n}?  In that case, we’ll have Φ(A,B)≤2log2|S|, but only for the boring reason that there’s hardly any entropy in B as a whole, to either influence A or be influenced by it.  For this reason, Tononi proposes a fix where we normalize each Φ(A,B) by dividing it by min{|A|,|B|}.  He then defines the integrated information Φ to be Φ(A,B), for whichever partition (A,B) minimizes the ratio Φ(A,B) / min{|A|,|B|}.  (Unless I missed it, Tononi never specifies what we should do if there are multiple (A,B)’s that all achieve the same minimum of Φ(A,B) / min{|A|,|B|}.  I’ll return to that point later, along with other idiosyncrasies of the normalization procedure.)

Tononi gives some simple examples of the computation of Φ, showing that it is indeed larger for systems that are more “richly interconnected” in an intuitive sense.  He speculates, plausibly, that Φ is quite large for (some reasonable model of) the interconnection network of the human brain—and probably larger for the brain than for typical electronic devices (which tend to be highly modular in design, thereby decreasing their Φ), or, let’s say, than for other organs like the pancreas.  Ambitiously, he even speculates at length about how a large value of Φ might be connected to the phenomenology of consciousness.

To be sure, empirical work in integrated information theory has been hampered by three difficulties.  The first difficulty is that we don’t know the detailed interconnection network of the human brain.  The second difficulty is that it’s not even clear what we should define that network to be: for example, as a crude first attempt, should we assign a Boolean variable to each neuron, which equals 1 if the neuron is currently firing and 0 if it’s not firing, and let f be the function that updates those variables over a timescale of, say, a millisecond?  What other variables do we need—firing rates, internal states of the neurons, neurotransmitter levels?  Is choosing many of these variables uniformly at random (for the purpose of calculating Φ) really a reasonable way to “randomize” the variables, and if not, what other prescription should we use?

The third and final difficulty is that, even if we knew exactly what we meant by “the f and x corresponding to the human brain,” and even if we had complete knowledge of that f and x, computing Φ(f,x) could still be computationally intractable.  For recall that the definition of Φ involved minimizing a quantity over all the exponentially-many possible bipartitions of {1,…,n}.  While it’s not directly relevant to my arguments in this post, I leave it as a challenge for interested readers to pin down the computational complexity of approximating Φ to some reasonable precision, assuming that f is specified by a polynomial-size Boolean circuit, or alternatively, by an NC0 function (i.e., a function each of whose outputs depends on only a constant number of the inputs).  (Presumably Φ will be #P-hard to calculate exactly, but only because calculating entropy exactly is a #P-hard problem—that’s not interesting.)

I conjecture that approximating Φ is an NP-hard problem, even for restricted families of f’s like NC0 circuits—which invites the amusing thought that God, or Nature, would need to solve an NP-hard problem just to decide whether or not to imbue a given physical system with consciousness!  (Alas, if you wanted to exploit this as a practical approach for solving NP-complete problems such as 3SAT, you’d need to do a rather drastic experiment on your own brain—an experiment whose result would be to render you unconscious if your 3SAT instance was satisfiable, or conscious if it was unsatisfiable!  In neither case would you be able to communicate the outcome of the experiment to anyone else, nor would you have any recollection of the outcome after the experiment was finished.)  In the other direction, it would also be interesting to upper-bound the complexity of approximating Φ.  Because of the need to estimate the entropies of distributions (even given a bipartition (A,B)), I don’t know that this problem is in NP—the best I can observe is that it’s in AM.

In any case, my own reason for rejecting IIT has nothing to do with any of the “merely practical” issues above: neither the difficulty of defining f and x, nor the difficulty of learning them, nor the difficulty of calculating Φ(f,x).  My reason is much more basic, striking directly at the hypothesized link between “integrated information” and consciousness.  Specifically, I claim the following:

Yes, it might be a decent rule of thumb that, if you want to know which brain regions (for example) are associated with consciousness, you should start by looking for regions with lots of information integration.  And yes, it’s even possible, for all I know, that having a large Φ-value is one necessary condition among many for a physical system to be conscious.  However, having a large Φ-value is certainly not a sufficient condition for consciousness, or even for the appearance of consciousness.  As a consequence, Φ can’t possibly capture the essence of what makes a physical system conscious, or even of what makes a system look conscious to external observers.

The demonstration of this claim is embarrassingly simple.  Let S=Fp, where p is some prime sufficiently larger than n, and let V be an n×n Vandermonde matrix over Fp—that is, a matrix whose (i,j) entry equals ij-1 (mod p).  Then let f:Sn→Sn be the update function defined by f(x)=Vx.  Now, for p large enough, the Vandermonde matrix is well-known to have the property that every submatrix is full-rank (i.e., “every submatrix preserves all the information that it’s possible to preserve about the part of x that it acts on”).  And this implies that, regardless of which bipartition (A,B) of {1,…,n} we choose, we’ll get

EI(A→B) = EI(B→A) = min{|A|,|B|} log2p,

and hence

Φ(A,B) = EI(A→B) + EI(B→A) = 2 min{|A|,|B|} log2p,

or after normalizing,

Φ(A,B) / min{|A|,|B|} = 2 log2p.

Or in words: the normalized information integration has the same value—namely, the maximum value!—for every possible bipartition.  Now, I’d like to proceed from here to a determination of Φ itself, but I’m prevented from doing so by the ambiguity in the definition of Φ that I noted earlier.  Namely, since every bipartition (A,B) minimizes the normalized value Φ(A,B) / min{|A|,|B|}, in theory I ought to be able to pick any of them for the purpose of calculating Φ.  But the unnormalized value Φ(A,B), which gives the final Φ, can vary greatly, across bipartitions: from 2 log2p (if min{|A|,|B|}=1) all the way up to n log2p (if min{|A|,|B|}=n/2).  So at this point, Φ is simply undefined.

On the other hand, I can solve this problem, and make Φ well-defined, by an ironic little hack.  The hack is to replace the Vandermonde matrix V by an n×n matrix W, which consists of the first n/2 rows of the Vandermonde matrix each repeated twice (assume for simplicity that n is a multiple of 4).  As before, we let f(x)=Wx.  Then if we set A={1,…,n/2} and B={n/2+1,…,n}, we can achieve

EI(A→B) = EI(B→A) = (n/4) log2p,

Φ(A,B) = EI(A→B) + EI(B→A) = (n/2) log2p,

and hence

Φ(A,B) / min{|A|,|B|} = log2p.

In this case, I claim that the above is the unique bipartition that minimizes the normalized integrated information Φ(A,B) / min{|A|,|B|}, up to trivial reorderings of the rows.  To prove this claim: if |A|=|B|=n/2, then clearly we minimize Φ(A,B) by maximizing the number of repeated rows in A and the number of repeated rows in B, exactly as we did above.  Thus, assume |A|≤|B| (the case |B|≤|A| is analogous).  Then clearly

EI(B→A) ≥ |A|/2,

while

EI(A→B) ≥ min{|A|, |B|/2}.

So if we let |A|=cn and |B|=(1-c)n for some c∈(0,1/2], then

Φ(A,B) ≥ [c/2 + min{c, (1-c)/2}] n,

and

Φ(A,B) / min{|A|,|B|} = Φ(A,B) / |A| = 1/2 + min{1, 1/(2c) – 1/2}.

But the above expression is uniquely minimized when c=1/2.  Hence the normalized integrated information is minimized essentially uniquely by setting A={1,…,n/2} and B={n/2+1,…,n}, and we get

Φ = Φ(A,B) = (n/2) log2p,

which is quite a large value (only a factor of 2 less than the trivial upper bound of n log2p).

Now, why did I call the switch from V to W an “ironic little hack”?  Because, in order to ensure a large value of Φ, I decreased—by a factor of 2, in fact—the amount of “information integration” that was intuitively happening in my system!  I did that in order to decrease the normalized value Φ(A,B) / min{|A|,|B|} for the particular bipartition (A,B) that I cared about, thereby ensuring that that (A,B) would be chosen over all the other bipartitions, thereby increasing the final, unnormalized value Φ(A,B) that Tononi’s prescription tells me to return.  I hope I’m not alone in fearing that this illustrates a disturbing non-robustness in the definition of Φ.

But let’s leave that issue aside; maybe it can be ameliorated by fiddling with the definition.  The broader point is this: I’ve shown that my system—the system that simply applies the matrix W to an input vector x—has an enormous amount of integrated information Φ.  Indeed, this system’s Φ equals half of its entire information content.  So for example, if n were 1014 or so—something that wouldn’t be hard to arrange with existing computers—then this system’s Φ would exceed any plausible upper bound on the integrated information content of the human brain.

And yet this Vandermonde system doesn’t even come close to doing anything that we’d want to call intelligent, let alone conscious!  When you apply the Vandermonde matrix to a vector, all you’re really doing is mapping the list of coefficients of a degree-(n-1) polynomial over Fp, to the values of the polynomial on the n points 0,1,…,n-1.  Now, evaluating a polynomial on a set of points turns out to be an excellent way to achieve “integrated information,” with every subset of outputs as correlated with every subset of inputs as it could possibly be.  In fact, that’s precisely why polynomials are used so heavily in error-correcting codes, such as the Reed-Solomon code, employed (among many other places) in CD’s and DVD’s.  But that doesn’t imply that every time you start up your DVD player you’re lighting the fire of consciousness.  It doesn’t even hint at such a thing.  All it tells us is that you can have integrated information without consciousness (or even intelligence)—just like you can have computation without consciousness, and unpredictability without consciousness, and electricity without consciousness.

It might be objected that, in defining my “Vandermonde system,” I was too abstract and mathematical.  I said that the system maps the input vector x to the output vector Wx, but I didn’t say anything about how it did so.  To perform a computation—even a computation as simple as a matrix-vector multiply—won’t we need a physical network of wires, logic gates, and so forth?  And in any realistic such network, won’t each logic gate be directly connected to at most a few other gates, rather than to billions of them?  And if we define the integrated information Φ, not directly in terms of the inputs and outputs of the function f(x)=Wx, but in terms of all the actual logic gates involved in computing f, isn’t it possible or even likely that Φ will go back down?

This is a good objection, but I don’t think it can rescue IIT.  For we can achieve the same qualitative effect that I illustrated with the Vandermonde matrix—the same “global information integration,” in which every large set of outputs depends heavily on every large set of inputs—even using much “sparser” computations, ones where each individual output depends on only a few of the inputs.  This is precisely the idea behind low-density parity check (LDPC) codes, which have had a major impact on coding theory over the past two decades.  Of course, one would need to muck around a bit to construct a physical system based on LDPC codes whose integrated information Φ was provably large, and for which there were no wildly-unbalanced bipartitions that achieved lower Φ(A,B)/min{|A|,|B|} values than the balanced bipartitions one cared about.  But I feel safe in asserting that this could be done, similarly to how I did it with the Vandermonde matrix.

More generally, we can achieve pretty good information integration by hooking together logic gates according to any bipartite expander graph: that is, any graph with n vertices on each side, such that every k vertices on the left side are connected to at least min{(1+ε)k,n} vertices on the right side, for some constant ε>0.  And it’s well-known how to create expander graphs whose degree (i.e., the number of edges incident to each vertex, or the number of wires coming out of each logic gate) is a constant, such as 3.  One can do so either by plunking down edges at random, or (less trivially) by explicit constructions from algebra or combinatorics.  And as indicated in the title of this post, I feel 100% confident in saying that the so-constructed expander graphs are not conscious!  The brain might be an expander, but not every expander is a brain.

Before winding down this post, I can’t resist telling you that the concept of integrated information (though it wasn’t called that) played an interesting role in computational complexity in the 1970s.  As I understand the history, Leslie Valiant conjectured that Boolean functions f:{0,1}n→{0,1}n with a high degree of “information integration” (such as discrete analogues of the Fourier transform) might be good candidates for proving circuit lower bounds, which in turn might be baby steps toward P≠NP.  More strongly, Valiant conjectured that the property of information integration, all by itself, implied that such functions had to be at least somewhat computationally complex—i.e., that they couldn’t be computed by circuits of size O(n), or even required circuits of size Ω(n log n).  Alas, that hope was refuted by Valiant’s later discovery of linear-size superconcentrators.  Just as information integration doesn’t suffice for intelligence or consciousness, so Valiant learned that information integration doesn’t suffice for circuit lower bounds either.

As humans, we seem to have the intuition that global integration of information is such a powerful property that no “simple” or “mundane” computational process could possibly achieve it.  But our intuition is wrong.  If it were right, then we wouldn’t have linear-size superconcentrators or LDPC codes.

I should mention that I had the privilege of briefly speaking with Giulio Tononi (as well as his collaborator, Christof Koch) this winter at an FQXi conference in Puerto Rico.  At that time, I challenged Tononi with a much cruder, handwavier version of some of the same points that I made above.  Tononi’s response, as best as I can reconstruct it, was that it’s wrong to approach IIT like a mathematician; instead one needs to start “from the inside,” with the phenomenology of consciousness, and only then try to build general theories that can be tested against counterexamples.  This response perplexed me: of course you can start from phenomenology, or from anything else you like, when constructing your theory of consciousness.  However, once your theory has been constructed, surely it’s then fair game for others to try to refute it with counterexamples?  And surely the theory should be judged, like anything else in science or philosophy, by how well it withstands such attacks?

But let me end on a positive note.  In my opinion, the fact that Integrated Information Theory is wrong—demonstrably wrong, for reasons that go to its core—puts it in something like the top 2% of all mathematical theories of consciousness ever proposed.  Almost all competing theories of consciousness, it seems to me, have been so vague, fluffy, and malleable that they can only aspire to wrongness.

[Endnote: See also this related post, by the philosopher Eric Schwetzgebel: Why Tononi Should Think That the United States Is Conscious.  While the discussion is much more informal, and the proposed counterexample more debatable, the basic objection to IIT is the same.]


Update (5/22): Here are a few clarifications of this post that might be helpful.

(1) The stuff about zombies and the Hard Problem was simply meant as motivation and background for what I called the “Pretty-Hard Problem of Consciousness”—the problem that I take IIT to be addressing.  You can disagree with the zombie stuff without it having any effect on my arguments about IIT.

(2) I wasn’t arguing in this post that dualism is true, or that consciousness is irreducibly mysterious, or that there could never be any convincing theory that told us how much consciousness was present in a physical system.  All I was arguing was that, at any rate, IIT is not such a theory.

(3) Yes, it’s true that my demonstration of IIT’s falsehood assumes—as an axiom, if you like—that while we might not know exactly what we mean by “consciousness,” at any rate we’re talking about something that humans have to a greater extent than DVD players.  If you reject that axiom, then I’d simply want to define a new word for a certain quality that non-anesthetized humans seem to have and that DVD players seem not to, and clarify that that other quality is the one I’m interested in.

(4) For my counterexample, the reason I chose the Vandermonde matrix is not merely that it’s invertible, but that all of its submatrices are full-rank.  This is the property that’s relevant for producing a large value of the integrated information Φ; by contrast, note that the identity matrix is invertible, but produces a system with Φ=0.  (As another note, if we work over a large enough field, then a random matrix will have this same property with high probability—but I wanted an explicit example, and while the Vandermonde is far from the only one, it’s one of the simplest.)

(5) The n×n Vandermonde matrix only does what I want if we work over (say) a prime field Fp with p>>n elements.  Thus, it’s natural to wonder whether similar examples exist where the basic system variables are bits, rather than elements of Fp.  The answer is yes. One way to get such examples is using the low-density parity check codes that I mention in the post.  Another common way to get Boolean examples, and which is also used in practice in error-correcting codes, is to start with the Vandermonde matrix (a.k.a. the Reed-Solomon code), and then combine it with an additional component that encodes the elements of Fp as strings of bits in some way.  Of course, you then need to check that doing this doesn’t harm the properties of the original Vandermonde matrix that you cared about (e.g., the “information integration”) too much, which causes some additional complication.

(6) Finally, it might be objected that my counterexamples ignored the issue of dynamics and “feedback loops”: they all consisted of unidirectional processes, which map inputs to outputs and then halt.  However, this can be fixed by the simple expedient of iterating the process over and over!  I.e., first map x to Wx, then map Wx to W2x, and so on.  The integrated information should then be the same as in the unidirectional case.


Update (5/24): See a very interesting comment by David Chalmers.

24 May 16:33

Functionally connected brain regions in the network activated during capsaicin inhalation

by Michael J. Farrell, Saskia Koch, Ayaka Ando, Leonie J. Cole, Gary F. Egan, Stuart B. Mazzone
Nosimpler

People get paid to do this.

Abstract

Coughing and the urge-to-cough are important mechanisms that protect the patency of the airways, and are coordinated by the brain. Inhaling a noxious substance leads to a widely distributed network of responses in the brain that are likely to reflect multiple functional processes requisite for perceiving, appraising, and behaviorally responding to airway challenge. The broader brain network responding to airway challenge likely contains subnetworks that are involved in the component functions required for coordinated protective behaviors. Functional connectivity analyses were used to determine whether brain responses to airway challenge could be differentiated regionally during inhalation of the tussive substance capsaicin. Seed regions were defined according to outcomes of previous activation studies that identified regional brain responses consistent with cough suppression, stimulus intensity coding, and perception of urge-to-cough. The subnetworks during continuous inhalation of capsaicin recapitulated the distributed regions previously implicated in discrete functional components of airway challenge. The outcomes of this study highlight the central representation of airways defence as a distributed network. Hum Brain Mapp 35:5341–5355, 2014. © 2014 Wiley Periodicals, Inc.

22 May 21:20

A Tutorial on Time-Evolving Dynamical Bayesian Inference. (arXiv:1305.0041v3 [physics.data-an] UPDATED)

by Tomislav Stankovski, Andrea Duggento, Peter V. E. McClintock, Aneta Stefanovska
Nosimpler

This might be Bayesian stuff I can actually get behind.

In view of the current availability and variety of measured data, there is an increasing demand for powerful signal processing tools that can cope successfully with the associated problems that often arise when data are being analysed. In practice many of the data-generating systems are not only time-variable, but also influenced by neighbouring systems and subject to random fluctuations (noise) from their environments. To encompass problems of this kind, we present a tutorial about the dynamical Bayesian inference of time-evolving coupled systems in the presence of noise. It includes the necessary theoretical description and the algorithms for its implementation. For general programming purposes, a pseudocode description is also given. Examples based on coupled phase and limit-cycle oscillators illustrate the salient features of phase dynamics inference. State domain inference is illustrated with an example of coupled chaotic oscillators. The applicability of the latter example to secure communications based on the modulation of coupling functions is outlined. MatLab codes for implementation of the method, as well as for the explicit examples, accompany the tutorial.

22 May 01:15

Jury Power Gets a Courtroom Nod in Possible Boost for Nullification

by J.D. Tuccille

JuryEarlier this month, the authority of the jury received a welcome nod from the Fifth Circuit Court of Appeals. In a case in which the defendant confessed in the courtroom to all of the charges against him, the court ruled that a directed "guilty" verdict was out of line, since the jury still had the right to make its own decision by its own criteria, no matter what the judge thought. For fans of jury nullification, here's a new endorsement of the power of a jury to bring a "not guilty" verdict for reasons of its own.

In the case of United State of America v. Juan Agudin Salazar, Judge Jerry E. Smith wrote:

Juan Salazar was charged with multiple drug and gun violations. At trial, the government presented overwhelming evidence of  guilt; against the advice of counsel, Salazar decided to testify and confessed to all of the crimes charged.  At the trial’s conclusion, believing no factual issue remained for the jury, the district court  instructed the jury "to go back and find the Defendant guilty." Because the Sixth Amendment safeguards even an obviously guilty defendant’s right to have a jury decide guilt or innocence, we vacate the conviction and remand.

Smith went on to point out that Salazar may have confessed, but he hadn't changed his plea. That left the ultimate decision of "guilty" or "not guilty" in the hands of the jury. Under the Sixth Amendment, he wrote, "a defendant's confession merely amounts to more, albeit compelling, evidence against him. But no amount of compelling evidence can override the right to have a jury determine his guilt."

The Fully Informed Jury Association suggests this case "has in effect re-affirmed the right of jury nullification, in which jurors may conscientiously deliver a Not Guilty verdict even in the face of overwhelming evidence that the law has technically been broken."

That's probably not what Judge Smith and his colleagues had in mind. But accidental victories can still be chalked up in the "win" column.

21 May 14:44

Oscillations: Synchrony shows mice the way

by Leonie Welberg
Nosimpler

Not sure if I have access to BU's library anymore, but this could be an interesting review...

Nature Reviews Neuroscience 15, 347 (2014). doi:10.1038/nrn3756

Author: Leonie Welberg

Synchronized, high-frequency gamma oscillations contribute to successful decision making in a working-memory task in mice and may, therefore, reflect content in working memory entering the animal's 'awareness'.

19 May 16:36

Nobody Could Have Predicted

by noreply@blogger.com (Atrios)
For awhile the big trend was international campuses, seen as a way to have a money spigot aimed at the home institution, or at least aimed at certain people in those institutions. Not everybody thought it was a good idea, of course.
Facing criticism for venturing into a country where dissent is not tolerated and labor can resemble indentured servitude, N.Y.U. in 2009 issued a “statement of labor values” that it said would guarantee fair treatment of workers. But interviews by The New York Times with dozens of workers who built N.Y.U.’s recently completed campus found that conditions on the project were often starkly different from the ideal.

Virtually every one said he had to pay recruitment fees of up to a year’s wages to get his job and had never been reimbursed. N.Y.U.’s list of labor values said that contractors are supposed to pay back all such fees. Most of the men described having to work 11 or 12 hours a day, six or seven days a week, just to earn close to what they had originally been promised, despite a provision in the labor statement that overtime should be voluntary.
17 May 18:11

Dan Piponi (sigfpe): Types, and two approaches to problem solving

by noreply@blogger.com (Dan Piponi)

Introduction

There are two broad approaches to problem solving that I see frequently in mathematics and computing. One is attacking a problem via subproblems, and another is attacking a problem via quotient problems. The former is well known though I’ll give some examples to make things clear. The latter can be harder to recognise but there is one example that just about everyone has known since infancy.


Subproblems

Consider sorting algorithms. A large class of sorting algorithms, including quicksort, break a sequence of values into two pieces. The two pieces are smaller so they are easier to sort. We sort those pieces and then combine them, using some kind of merge operation, to give an ordered version of the original sequence. Breaking things down into subproblems is ubiquitous and is useful far outside of mathematics and computing: in cooking, in finding our path from A to B, in learning the contents of a book. So I don’t need to say much more here.


Quotient problems

The term quotient is a technical term from mathematics. But I want to use the term loosely to mean something like this: a quotient problem is what a problem looks like if you wear a certain kind of filter over your eyes. The filter hides some aspect of the problem that simplifies it. You solve the simplified problem and then take off the filter. You now ‘lift’ the solution of the simplified problem to a solution to the full problem. The catch is that your filter needs to match your problem so I’ll start by giving an example where the filter doesn’t work.


Suppose we want to add a list of integers, say: 123, 423, 934, 114. We can try simplifying this problem by wearing a filter that makes numbers fuzzy so we can’t distinguish numbers that differ by less than 10. When we wear this filter 123 looks like 120, 423 looks like 420, 934 looks like 930 and 114 looks like 110. So we can try adding 120+420+930+110. This is a simplified problem and in fact this is a common technique to get approximate answers via mental arithmetic. We get 1580. We might hope that when wearing our filters, 1580 looks like the correct answer. But it doesn’t. The correct answer is 1594. This filter doesn’t respect addition in the sense that if a looks like a’ and b looks like b’ it doesn’t follow that a+b looks like a’+b’.


To solve a problem via quotient problems we usually need to find a filter that does respect the original problem. So let’s wear a different filter that allows us just to see the last digit of a number. Our original problem now looks like summing the list 3, 3, 4, 4. We get 4. This is the correct last digit. If we now try a filter that allows us to see just the last two digits we see that summing 23, 23, 34, 14 does in fact give the correct last two digits. This is why the standard elementary school algorithms for addition and multiplication work through the digits from right to left: at each stage we’re solving a quotient problem but the filter only respects the original problem if it allows us to see the digits to the right of some point, not digits to the left. This filter does respect addition in the sense that if a looks like a’ and b looks like b’ then a+b looks like a’+b’.


Another example of the quotient approach is to look at the knight’s tour problem in the case where two opposite corners have been removed from the chessboard. A knight’s tour is a sequence of knight’s moves that visit each square on a board exactly once. If we remove opposite corners of the chessboard, there is no knight’s tour of the remaining 62 squares. How can we prove this? If you don’t see the trick you can get get caught up in all kinds of complicated reasoning. So now put on a filter that removes your ability to see the spatial relationships between the squares so you can only see the colours of the squares. This respects the original problem in the sense that a knight’s move goes from a black square to a white square, or from a white square to a black square. The filter doesn’t stop us seeing this. But now it’s easier to see that there are two more squares of one colour than the other and so no knight’s tour is possible. We didn’t need to be able to see the spatial relationships at all.


(Note that this is the same trick as we use for arithmetic, though it’s not immediately obvious. If we think of the spatial position of a square as being given by a pair of integers (x, y), then the colour is given by x+y modulo 2. In other words, by the last digit of x+y written in binary. So it’s just the see-only-digits-on-the-right filter at work again.)


Wearing filters while programming

So now think about developing some code in a dynamic language like Python. Suppose we execute the line:


a = 1


The Python interpreter doesn’t just store the integer 1 somewhere in memory. It also stores a tag indicating that the data is to be interpreted as an integer. When you come to execute the line:


b = a+1


it will first examine the tag in a indicating its type, in this case int, and use that to determine what the type for b should be.


Now suppose we wear a filter that allows us to see the tag indicating the type of some data, but not the data itself. Can we still reason about what our program does?


In many cases we can. For example we can, in principle, deduce the type of


a+b*(c+1)/(2+d)


if we know the types of a, b, c, d. (As I’ve said once before, it’s hard to make any reliable statement about a bit of Python code so let's suppose that a, b, c and d are all either of type int or type float.) We can read and understand quite a bit of Python code wearing this filter. But it’s easy to go wrong. For example consider


if a>1 then:
return 1.0
else:
return 1


The type of the result depends on the value of the variable a. So if we’re wearing the filter that hides the data, then we can’t predict what this snippet of code does. When we run it, it might return an int sometimes and a float other times, and we won’t be able to see what made the difference.


In a statically typed language you can predict the type of an expression knowing the type of its parts. This means you can reason reliably about code while wearing the hide-the-value filter. This means that almost any programming problem can be split into two parts: a quotient problem where you forget about the values, and then problem of lifting a solution to the quotient problem to a solution to the full problem. Or to put that in more conventional language: designing your data and function types, and then implementing the code that fits those types.


I chose to make the contrast between dynamic and static languages just to make the ideas clear but actually you can happily use similar reasoning for both types of language. Compilers for statically typed languages, give you a lot of assistance if you choose to solve your programming problems this way.


A good example of this at work is given in Haskell. If you're writing a compiler, say, you might want to represent a piece of code as an abstract syntax tree, and implement algorithms that recurse through the tree. In Haskell the type system is strong enough that once you’ve defined the tree type the form of the recursion algorithms is often more or less given. In fact, it can be tricky to implement tree recursion incorrectly and have the code compile without errors. Solving the quotient problem of getting the types right gets you much of the way towards solving the full problem.


And that’s my main point: types aren’t simply a restriction mechanism to help you avoid making mistakes. Instead they are a way to reduce some complex programming problems to simpler ones. But the simpler problem isn’t a subproblem, it’s a quotient problem.

Dependent types

Dependently typed languages give you even more flexibility with what filters you wear. They allow you to mix up values and types. For example both C++ and Agda (to pick an unlikely pair) allow you to wear filters that hide the values of elements in your arrays while allowing you to see the length of your arrays. This makes it easier to concentrate on some aspects of your problem while completely ignoring others.


Notes

I wrote the first draft of this a couple of years ago but never published it. I was motivated to post by a discussion kicked off by Voevodsky on the TYPES mailing list http://lists.seas.upenn.edu/pipermail/types-list/2014/001745.html


This article isn’t a piece of rigorous mathematics and I’m using mathematical terms as analogies.


The notion of a subproblem isn’t completely distinct from a quotient problem. Some problems are both, and in fact some problems can be solved by transforming them so they become both.

More generally, looking at computer programs through different filters is one approach to abstract interpretation http://en.wikipedia.org/wiki/Abstract_interpretation. The intuition section there (http://en.wikipedia.org/wiki/Abstract_interpretation#Intuition) has much in common with what I’m saying.
17 May 18:09

Assistant Police Chief Says "Too Many Sacrifices" to Allow Flag Defacing In His Pennsylvania Community

by Ed Krayewski
Nosimpler

Haha remember the first Elk House Halloween party?

trigger warningThe U.S. Code makes desecration of an American flag a crime punishable by up to one year in jail. Most states have their own additional laws. In Pennsylvania it is apparently illegal to "disrespect" the U.S. flag or to use it for publicity or commercial purposes. (No one tell all the businesses in Philly, and elsewhere I'm sure, that use American flags to advertise!)

Joshua Brubacker, of Alleghany Township, found out about laws protecting the American flag last week after spray painting his and flying it upside down in protest, he says, of the sale of Wounded Knee. Via WPXI:

"I was offended by it when I first saw it. I had an individual stop here at the station, a female, who was in the military and she was very offended by it." said Allegheny Township police Assistant Chief L.J. Berg.

Berg said he took the flag down and charged Joshua Brubaker with desecration and insults to the American flag….

"People have made too many sacrifices to protect the flag and to have this happen in my community, I'm not happy with that," Berg said.

Brubaker said he wishes those who were offended would have come to him so he could have explained his position. He is facing misdemeanor charges, but hopes police will reconsider. 

Brubaker also claims he never meant to offend someone which begs the question of what he did intend with displaying a defaced, upside down flag? Protest isn't all that effective if it doesn't upset somebody—it's kind of the point. Brubaker might find more success appealing to his First Amendment rights rather than playing along with (or buying into, as the case may be) the ridiculous notion that in America offensiveness is something the police should, uh, police, or that the right not to be offended is why Americans make sacrifices for their country.

h/t Anthony B.

14 May 23:07

The Problem of Military Robot Ethics

by Peter Suderman
Nosimpler

I would make an exception to my military funding ban for this.

I’m not entirely sure what it would mean for a robot to have morals, but the U.S. military is about to spend $7.5 million to try to find out. As J.D. Tuccille noted earlier today, the Office of Naval Research has awarded grants to artificial intelligence (A.I.) researchers at multiple universities to “explore how to build a sense of right and wrong and moral consequence into autonomous robotic systems,” reports Defense One.

The science-fiction-friendly problems with creating moral robots get ponderous pretty fast, especially when the military is involved: What sorts of ethical judgments should a robot make? How to prioritize between two competing moral claims when conflict inevitably arises? Do we define moral and ethical judgments as somehow outside the realm of logic—and if so, how does a machine built on logical operations make those sorts of considerations? I could go on.

You could perhaps head off a lot of potential problems by installing behavioral restrictions along the lines of Isaac Asimov’s Three Laws of Robotics, which state that robots can’t harm people or even allow harm through inaction, must obey people unless it could cause someone harm, and must protect themselves, except when that conflicts with the other two laws. But in a military context, where robots would at least be aiding with a war effort, even if only in a secondary capacity, those sorts of no-harm-to-humans rules would probably prove unworkable.

That’s especially true if the military ends up pursuing autonomous fighting machines, what most people would probably refer to as killer robots. As the Defense One story notes, the military currently prohibits fully autonomous machines from using lethal force, and even semi-autonomous drones or others are not allowed to select and engage targets without prior authorization from a human. But one U.N. human rights official warned last year that it’s likely that governments will eventually begin creating fully autonomous lethal machines. Killer robots! Coming soon to a government near you.

Obviously Asimov’s Three Laws wouldn’t work on a machine designed to kill. Would any moral or ethical system? It seems plausible that you could build in rules that work basically like the safety functions of many machines today, in which the specific conditions result in safety behaviors or shut down orders. But it’s hard to imagine, say, an attack drone with an ethical system that allows it to make decisions about right and wrong in a battlefield context.

What would that even look like? Programming problems aside, the moral calculus involved in waing war is too murky and too widely disputed to install in a machine. You can’t even get people to come to any sort of agreement on the morality of using drones for targeted killing today, when they are almost entirely human controlled. An artificial intelligence designed to do the same thing would just muddy the moral waters even further.

Indeed, it’s hard to imagine even a non-lethal military robot with a meaningful moral mental system, especially if we’re pushing into the realm of artificial intelligence. War always presents ethical problems, and no software system is likely to resolve them. If we end up building true A.I. robots, then, we’ll probably just have to let them decide for themselves. 

14 May 04:21

Circenses without panem isn't sufficient

by Minnesotastan

An article at AlterNet notes that if you don't appease the public with both bread and circuses*, social unrest may develop.
From 2008 to 2014, insurrectionist activity has sequentially erupted across the globe, from Tunisia and Egypt to Syria and Yemen; from Greece, Spain, Turkey and Brazil to Thailand, Bosnia, Venezuela and the Ukraine...

...beneath what we’ve come to perceive as isolated and distinct events is a shared but neglected root cause of environmental crisis. What most people don’t realize is that outbreaks of social unrest are preceded, usually, by a single pattern — an unholy trinity of drought, low crop yield and soaring food prices.

So what do the Arab Spring, Syrian civil war, Occupy Gezi, and the recent conflicts in the Ukraine, Venezuela, Bosnia and Thailand all have in common? Expensive food… and not much of it.
Details at the link.

* "This phrase originates from Rome in Satire X of the Roman satirist and poet Juvenal (circa A.D. 100). In context, the Latin metaphor panem et circenses (bread and circuses) identifies the only remaining cares of a new Roman populace which cares not for its historical birthright of political involvement. Here Juvenal displays his contempt for the declining heroism of his contemporary Romans. Roman politicians devised a plan in 140 B.C. to win the votes of these new citizens: giving out cheap food and entertainment, "bread and circuses", would be the most effective way to rise to power."
14 May 04:12

How multiplicity determines entropy [Physics]

by Hanel, R., Thurner, S., Gell-Mann, M.
The maximum entropy principle (MEP) is a method for obtaining the most likely distribution functions of observables from statistical systems by maximizing entropy under constraints. The MEP has found hundreds of applications in ergodic and Markovian systems in statistical mechanics, information theory, and statistics. For several decades there has been...
14 May 00:43

One-bit compressive sensing with norm estimation

by Igor
Laurent Jacques mentioned it on on his twitter feed but I think it is bigger than even what he says (more on that at the end of this entry):
Norm can be recovered! "One-bit compressive sensing with norm estimation" by K. Knudson, R. Saab, R. Ward http://t.co/bNnIVnE7jM #1bitcs
— Laurent Jacques (@jacquesdurden) May 3, 2014




Compressive sensing typically involves the recovery of a sparse signal x from linear measurements ai.x, but recent research has demonstrated that recovery is also possible when the observed measurements are quantized to sign(ai.x)∈{±1}. Since such measurements give no information on the norm of x, recovery methods geared towards such 1-bit compressive sensing measurements typically assume that ∥x∥2=1. We show that if one allows more generally for quantized affine measurements of the form sign(ai.x+bi), and if such measurements are random, an appropriate choice of the affine shifts bi allows norm recovery to be easily incorporated into existing methods for 1-bit compressive sensing. Alternatively, if one is interested in norm estimation alone, we show that the fraction of measurements quantized to +1 (versus −1) can be used to estimate ∥x∥2 through a single evaluation of the inverse Gaussian error function, providing a computationally efficient method for norm estimation from 1-bit compressive measurements. Finally, all of our recovery guarantees are universal over sparse vectors, in the sense that with high probability, one set of measurements will successfully estimate all sparse vectors simultaneously.


Like Laurent says, it is a big deal for the 1bit compressive sensing and quantization crowd but I want to take this further. 1-bit compressive sensing is first and foremost a nonlinear encoding of measurements of a special kind, in fact it very much parallels the first layer of neural networks except that it doesn't use a sigmoid and does not use an affine function. The paper lifts that problem by adding an afine constant and thereby shows that more of the signal can be recovered than initially thought. In short, it provides a natural connection between the whole theoretical framework of compressive sensing and the vibrant and exciting development in neural networks, see the recent:
In other words, some theoretical justification is coming to the neural network community and by the same token, this might provide a way to design those in a more efficient manner. Woohoo !


Join the CompressiveSensing subreddit or the Google+ Community and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.
13 May 13:50

Quiver Grassmannians can be anything

by lievenlb
Nosimpler

Yohan, to answer your question from Facebook about a month ago, the Grassmannian is a manifold in which subspaces are viewed as points. This all relates to the amplituhedron that got so much attention in theoretical physics last year (through the so-called positive Grassmannian), as well as various other hedra (associahedra, permutahedra, etc.). It's kinda funny to see quantum mechanics, abstract mathematics, and cognitive science (Tononi's notion of qualia spaces) essentially converge on objects that are basically gussied up versions of Platonic solids.

A standard Grassmannian $Gr(m,V)$ is the manifold having as its points all possible $m$-dimensional subspaces of a given vectorspace $V$. As an example, $Gr(1,V)$ is the set of lines through the origin in $V$ and therefore is the projective space $\mathbb{P}(V)$. Grassmannians are among the nicest projective varieties, they are smooth and allow a cell decomposition.

A quiver $Q$ is just an oriented graph. Here’s an example

Q~:~\xymatrix{\bullet & \bullet \ar[l]_a \ar@/^2ex/[r]^{b} \ar@/_2ex/[r]_{c} & \bullet}

A representation $V$ of a quiver assigns a vector-space to each vertex and a linear map between these vertex-spaces to every arrow. As an example, a representation $V$ of the quiver $Q$ consists of a triple of vector-spaces $(V_1,V_2,V_3)$ together with linear maps $f_a~:~V_2 \rightarrow V_1$ and $f_b,f_c~:~V_2 \rightarrow V_3$.

A sub-representation $W \subset V$ consists of subspaces of the vertex-spaces of $V$ and linear maps between them compatible with the maps of $V$. The dimension-vector of $W$ is the vector with components the dimensions of the vertex-spaces of $W$.

This means in the example that we require $f_a(W_2) \subset W_1$ and $f_b(W_2)$ and $f_c(W_2)$ to be subspaces of $W_3$. If the dimension of $W_i$ is $m_i$ then $m=(m_1,m_2,m_3)$ is the dimension vector of $W$.

The quiver-analogon of the Grassmannian $Gr(m,V)$ is the Quiver Grassmannian $QGr(m,V)$ where $V$ is a quiver-representation and $QGr(m,V)$ is the collection of all possible sub-representations $W \subset V$ with fixed dimension-vector $m$. One might expect these quiver Grassmannians to be rather nice projective varieties.

However, last week Markus Reineke posted a 2-page note on the arXiv proving that every projective variety is a quiver Grassmannian.

Let’s illustrate the argument by finding a quiver Grassmannian $QGr(m,V)$ isomorphic to the elliptic curve in $\mathbb{P}^2$ with homogeneous equation $Y^2Z=X^3+Z^3$.

Consider the Veronese embedding $\mathbb{P}^2 \hookrightarrow \mathbb{P}^9$ obtained by sending a point $(x:y:z)$ to the point

\[ (x^3:x^2y:x^2z:xy^2:xyz:xz^2:y^3:y^2z:yz^2:z^3) \]

The upshot being that the elliptic curve is now realized as the intersection of the image of $\mathbb{P}^2$ with the hyper-plane $\mathbb{V}(X_0-X_7+X_9)$ in the standard projective coordinates $(x_0:x_1:\cdots:x_9)$ for $\mathbb{P}^9$.

To describe the equations of the image of $\mathbb{P}^2$ in $\mathbb{P}^9$ consider the $6 \times 3$ matrix with the rows corresponding to $(x^2,xy,xz,y^2,yz,z^2)$ and the columns to $(x,y,z)$ and the entries being the multiplications, that is

$$\begin{bmatrix} x^3 & x^2y & x^2z \\ x^2y & xy^2 & xyz \\ x^2z & xyz & xz^2 \\ xy^2 & y^3 & y^2z \\ xyz & y^2z & yz^2 \\ xz^2 & yz^2 & z^3 \end{bmatrix} = \begin{bmatrix} x_0 & x_1 & x_2 \\ x_1 & x_3 & x_4 \\ x_2 & x_4 & x_5 \\ x_3 & x_6 & x_7 \\ x_4 & x_7 & x_8 \\ x_5 & x_8 & x_9 \end{bmatrix}$$

But then, a point $(x_0:x_1: \cdots : x_9)$ belongs to the image of $\mathbb{P}^2$ if (and only if) the matrix on the right-hand side has rank $1$ (that is, all its $2 \times 2$ minors vanish). Next, consider the quiver


\xymatrix{\bullet & & \bullet \ar[ll]^h \ar@/^2ex/[rr]^x \ar[rr]^y \ar@/_2ex/[rr]^z & & \bullet}

and consider the representation $V=(V_1,V_2,V_3)$ with vertex-spaces $V_1=\mathbb{C}$, $V_2 = \mathbb{C}^{10}$ and $V_2 = \mathbb{C}^6$. The linear maps $x,y$ and $z$ correspond to the columns of the matrix above, that is

$$(x_0,x_1,x_2,x_3,x_4,x_5,x_6,x_7,x_8,x_9) \begin{cases} \rightarrow^x~(x_0,x_1,x_2,x_3,x_4,x_5) \\ \rightarrow^y~(x_1,x_3,x_4,x_6,x_7,x_8) \\ \rightarrow^z~(x_2,x_4,x_5,x_7,x_8,x_9) \end{cases}$$

The linear map $h~:~\mathbb{C}^{10} \rightarrow \mathbb{C}$ encodes the equation of the hyper-plane, that is $h=x_0-x_7+x_9$.

Now consider the quiver Grassmannian $QGr(m,V)$ for the dimension vector $m=(0,1,1)$. A base-vector $p=(x_0,\cdots,x_9)$ of $W_2 = \mathbb{C}p$ of a subrepresentation $W=(0,W_2,W_3) \subset V$ must be such that $h(x)=0$, that is, $p$ determines a point of the hyper-plane.

Likewise the vectors $x(p),y(p)$ and $z(p)$ must all lie in the one-dimensional space $W_3 = \mathbb{C}$, that is, the right-hand side matrix above must have rank one and hence $p$ is a point in the image of $\mathbb{P}^2$ under the Veronese.

That is, $Gr(m,V)$ is isomorphic to the intersection of this image with the hyper-plane and hence is isomorphic to the elliptic curve.

The general case is similar as one can view any projective subvariety $X \hookrightarrow \mathbb{P}^n$ as isomorphic to the intersection of the image of a specific $d$-uple Veronese embedding $\mathbb{P}^n \hookrightarrow \mathbb{P}^N$ with a number of hyper-planes in $\mathbb{P}^N$.

13 May 13:24

Compressive Sensing and Short Term Memory / Visual Nonclassical Receptive Field Effects

by Igor


This is fantastic: Here are papers trying to link Short Term Memory with either the RIP or the Donoho-Tanner phase transition so that we have actual limits on what constitute a good neural assembly. You probably recall Chris Rozell is one the lead behind one of the rapid solver construction (see Faster Than a Blink of an Eye. ) Being fast is one thing, but we also want to see how that approach translates into system size and so forth (a little bit like what the GWAS folks are trying to do with the sample size of the genome using compressive sensing ( see Application of compressed sensing to genome wide association studies and genomic selection, Predicting the Future: The Upcoming Stephanie Events). Anyway, here is Chris just sent me :
:

Hi Igor-
I know you've had a long-standing interest in the intersection of neuroscience with sparsity (and possibly compressed sensing). I wanted to draw your attention to some recent papers at this intersection. 
First, you may find our recent paper in PLoS Computational Biology interesting: M. Zhu and C.J. Rozell. Visual nonclassical receptive field effects emerge from sparse coding in a dynamical system. PLoS Computational Biology, 9(8):e1003191, August 2013.
http://www.ploscompbiol.org/article/info:doi/10.1371/journal.pcbi.1003191

While the classic Olshausen & Field (1996) paper showed that sparse coding can account for classical receptive fields shapes in primary visual cortex (V1), that paper didn't say anything about whether sparse coding can actually explain response properties observed in V1 neurons. In fact, classical receptive fields are not very good predictors of V1 neural responses to natural scenes (especially single-trial responses). The paper above performs simulated electrophysiology experiments to demonstrate that a wide variety of observed nonlinear/nonclassical response properties (single cell and population) are emergent behaviors of a sparse coding model implemented in a dynamical system. These results show that the sparse coding hypothesis, when coupled with a biophysically plausible implementation, can provide a unified high-level functional interpretation to many response properties that have generally been viewed through distinct mechanistic or phenomenological models. 

Second, it's possible you saw a preprint version of this, but you may be interested in our recent paper that just appeared in Neural Computation: A.S. Charles, H.L. Yap, and C.J. Rozell. Short term memory capacity in networks via the restricted isometry property. Neural Computation, 26(6):1198-1235, June 2014. http://arxiv.org/abs/1307.7970 

An interesting open question is how brains can store sequence memories for lengths of time on the order of seconds, which is probably too short to be due to plasticity (changes in synaptic connections between cells). The conjecture is that this type of memory must be due to transient activity in recurrently connected networks. In addition to questions about biological memory, this question has also come up as part in trying to understand why reservoir computing strategies (i.e., untrained, randomly connected artificial neural networks such as Echo State Networks and Liquid State Machines) work so well in some situations. The conventional wisdom is that signal structure could be used by the network to dramatically increase memory capacity, but this was not supported by formal analysis. In the paper above, we use a compressed sensing style analysis to show conclusively that memory capacities for randomly connected networks can be much higher than what was previously known when the signal has sparse structure than can be exploited. From a technical perspective, this paper establishes RIP for a very structured random matrix that corresponds to the propagation of a signal through a networked system. 
regards,
chris rozell
Thanks Chris, this is outstanding. I will come back to this later.

Very relevant set of links:


Join the CompressiveSensing subreddit or the Google+ Community and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.
13 May 04:44

How Black People See White Culture - YouTube

09 May 18:24

"She flung her anus high up in the air..."

by Minnesotastan
A passage from the book Matisse on the Loose, as translated by an optical character reader:
"When she spotted me, she flung her anus high in the air and kept them up until she reached me. 'Matisse. Oh boy!' she said. She grabbed my anus and positioned my body in the direction of the east gallery and we started walking."
It's all explained Sarah Wendell, editor of Smart Bitches, Trashy Books:
"So if the text is old, and it says 'arms', the OCR [optical character recognition] scanner will see it as 'anus.' OMG," Wendell tweeted. (She was referring to optical character recognition, the process by which printed texts can be scanned and converted into ebooks.)
Via The Guardian and Melville House, where there are additional examples of the error.
07 May 16:47

"No copula, no problem"

by Minnesotastan
This is a followup to yesterday's post about crash blossoms, in the comments to which an anonymous reader clued me in to the term "zero copula":
Zero copula is a linguistic phenomenon whereby the subject is joined to the predicate without overt marking of this relationship (like the copula 'to be' in English)... used most frequently in rhetoric and casual speech...

Standard English exhibits a very limited form of the zero copula, common in statements like "The higher, the better"; "The more, the merrier"...

Zero copula also appears in casual questions and statements like "You from out of town?"; "Enough already!" where the verb (and more) may be omitted due to syncope. Apart from syncope, the zero copula is probably not used productively in standard English.

The zero copula is far more productive in Caribbean creoles and African American Vernacular English, some varieties of which regularly omit the copula. For instance, "You crazy!", "Where you at?" and "Who she?" As in Russian, this is the case only in the present tense. In past-tense sentences, the copula must be specified...

The zero copula is also present, in a slightly different and more regular form, in the headlines of English newspapers, where short words and articles are generally omitted to conserve space. For example, a headline would more likely say "Gulf coast in ruins" than "Gulf coast is in ruins"... 
The Wikipedia page goes on to discuss zero copula in languages other than English (note zero copula is standard in American Sign Language).
07 May 00:56

Pro Wrestling: Vladimir Putin vs. Edward Snowden

by Zenon Evans

Professional wrestling, with its monstrous egos, blowhard rhetoric, and bad solutions (use the chair!), is kind of like politics except that it's got better acting. And, you can actually use wrestling as a sort of barometer for the average person's views, instead of the government's, on hot-button issues. This past Sunday World Wrestling Entertainment (WWE) showed us how Americans perceive Russian President Vladimir Putin and whistle-blower Edward Snowden.

At a pay-per-view event in New Jersey, C.J. "Lana The Ravishing Russian" Perry riled up fans for a fight featuring Miroslav "Alexander Rusev" Barnyashev, by dedicating his performance to Putin. Perry announced in her best fake accent:

I am proud to be Russian. I am proud to come from a country with the most dominant and powerful president in the world, Vladimir Putin. He makes fools out of every one of you. You are merely pawns in his game of global dominance.

The jumbotron displayed Putin's mug. The crowd started booing and chanting U-S-A so loudly, you could feel the Cold War rekindling. The negative response to Putin was expected, since the WWE knows its marketing. It is, after all, a billion dollar business with 15 million weekly viewers. Unlike other sports that occasionally stumble into American's ongoing debates on race or drugs or sexuality, pro wrestling's scripted game explicitly relies on popular culture outside the sport and the ability to reflect them in a way the audience wants.

Still, the fact that the WWE is reviving nationalistic gimmicks not seen since Hulk Hogan spat on the Soviet flag, Rocky clocked Drago, and the Wolverines valiantly battled the Red Army, shows just how bad the perception of America's foreign relations have gotten lately. And it's not because there's a threat to American lives any more real than pro wrestling itself. Rather, it boils down to the fact that the Obama and Putin administrations have dropped the ball on two decades worth of relationship repairing the U.S. and Russia did after the Cold War ended.

But, that's not even where The Ravishing Russian's speech ended. Her final blow was supposed to be that Putin "welcomes with open arms the patriot Edward Snowden" in his quest for power.

The Washington Free Beacon, whose editors have collectively stated their dislike for Snowden and Russia, saw the match and reported "audible ire from the audience" for Snowden. You can judge for yourself, but that's not what I heard. The response from the crowd sounded conspicuously muted to me. The fans were worked up into a lather by the Putin bit but the Snowden insult, which was supposed to be the final straw, just fell flat.

It fell flat in the same way big government advocates and apologists like President Obama, Sen. John McCain (R-Ariz.), and Rep. Mike Rogers (R-Mich.) do every time they try to unambiguously paint Snowden as a traitor who deserves time in prison for drawing attention to the massive surveillance state that continues to violate virtually every American's privacy.

Wrestling and politics both rely on framing issues as dualities: red versus blue, face versus heel, with us or against us. But neither of these groups can elicit outrage about Snowden, because his personal legacy is too mixed. In a recent YouGov poll that asked if Snowden did "the right thing" by exposing government surveillance, an almost equal number of people said yes as said no, and slightly more said they weren't sure.

And, far less ambiguous are the feelings for government snoops that Snowden awakened in the public. Exposing the National Security Administration (NSA) immediately and dramatically shifted Americans' concerns from terrorism to civil liberty violations as one poll shows. Another survey demonstrates that a record number of Americans see big government as the greatest threat to the future of America. Reason-Rupe poll data indicates that only 18 percent of Americans trust the NSA with their personal information and that many people think the agency is violating their privacy.

The WWE is savvy. It pays attention to how fans feel and responds quickly and accordingly. The Russophobia will continue as long as the international political showboating does, but I bet they'll stop lumping Snowden in the bad guy camp after this goof. Public servants who value their own popularity should take note.

06 May 19:22

Nobody Has Any Money For Retirement

by noreply@blogger.com (Atrios)
Something's gotta change.
The median size of a 401(k) is $24,400 as of March 31, with people older than 55 having $65,300, according to Fidelity Investments. Those funds can disappear quickly in retirement, and the early withdrawals indicate that the coming retirement crisis could be even more acute than expected.