Thursday, November 22, 2018

Every odd number is the difference between two squares

From here via SSC. This is apparently a 'twitter' proof (ie short).



Algebraically (n + 1)2 - n2 = 2n + 1 which is odd.

A slight defence of homo economicus

Amazon link

I've just started this and on page 5 I read:
" Many times the information required to arrive at the theoretical optimum is not available to those who make the decisions, and even if the information is available, the people who make the decisions often do not have an incentive to implement the optimal policy."
If we could only be more like homo economicus, how much better capitalism would work. Cue the application of behavioural genetics in engineering mode.

---

Homo economicus is "... a portrayal of humans as agents who are consistently rational and narrowly self-interested, and who usually pursue their subjectively-defined ends optimally." (Wikipedia).

In fact public choice theory does assume homo economicus, but asserts that incentives often can't or don't align - eg the principal-agent problem, market failure, etc. In capitalist economies consisting of real people, things work even less well.

Who or what should we blame?

Wednesday, November 21, 2018

"Blueprint" - Robert Plomin (a review)

Amazon link

Plomin's book is written in a practical, down to earth style. It's not at all academic. The big ideas (relentlessly hammered home) are these.

  • All traits (physical and psychological) have a substantial genetic component. Heritability is around 50% as a rule of thumb. Higher for traits like height, weight and intelligence.

  • Shared environments such as the family and school do not significantly affect traits. Schools don't make you smarter and parents don't make you nicer.

  • Schools will educate you to your potential aptitude if they're any good (and may socialise you into an elite). Bad parenting can damage a child. But absent active damage, it's the child's inherited genetics which determine performance and personal outcomes.

  • Non-genetic influences on traits are the generally random effects of life events and short lasting (these include test imprecision).

  • Apparent environmental differences (eg parents who read to their children in a house full of books as against ...) have a substantial genetic component (studious parents have studious children). Plomin describes this as nature in nurture.

Strong claims, but based on vast experiments with hundreds of thousands of participants. The results strongly replicate. Problem is, they're quite counterintuitive and few people believe them.

Toby Young writes in Quillette about a recent debate in London:
"On Monday in London’s Emmanuel Centre a debate took place that pitted two Quillette contributors - Robert Plomin and Stuart Ritchie - against two “experts” on child psychology - Susan Pawlby and Ann Pleshette Murphy. The motion was “Parenting doesn’t matter (or not as much as you think)” ...

The ushers asked people to vote for or against the motion on their way in and then again at the end, the idea being that the “winners” would be the side that persuaded the most people to change their minds rather than the side that got the most votes. Which was just as well for Plomin and Ritchie since only 17 percent agreed with them at the beginning of the evening, with 66 percent against and 17 percent saying “Don’t Know.”
Young describes the debate in detail and then describes the results:
"The debate was well-chaired by Xand van Tulleken, a doctor and broadcaster who has an identical twin brother named Chris, and, after he’d taken plenty of questions and done his best to sum up, the audience was asked to vote again.

As expected, a majority still disagreed with the motion, but Plomin and Ritchie had succeeded in persuading some people to change their minds. The number against the motion declined from 66 percent to 51 percent, while those in favor increased from 17 percent to 29 percent, with 20 percent saying “Don’t Know.”

That made Plomin and Ritchie the winners."
Few people truly believe scientific abstractions until the engineering stares them in the face. But that could happen, given the falling cost of whole genome sequencing together with ever larger scale Genome-Wide Association Studies (GWAS) for every trait imaginable.

Soon anyone's genome will be cheaply sequenced and their polygenic scores read off (at any age including infancy, in utero or for IVF selection) to deliver personalised results for a palette of physical, psychological and aptitude life-traits.

Already there are the early signs that the liberal media, the op-ed writers and professional pundits are beginning to warm to the idea. Plomin seems to have avoided the public evisceration he was undoubtedly fearing.

The book is interesting, important and enlightening. One of the books of the year.

---

Eric Turkheimer on his blog is complaining that Plomin has used his ideas without attribution. This may be true but he should get over it: this is popularisation, not academia. If the reception had been super-hostile, Turkheimer might now be experiencing relief rather than angst.

Turkheimer is not alone, by the way. Few of Plomin's peers get namechecked. There could be a few more bruised egos out there.

---

James Thompson of UCL has an excellent and detailed chapter-by-chapter review and Greg Cochran wrote a review at Quillette, "Forget Nature Versus Nurture. Nature Has Won".

Worth checking both out.

Tuesday, November 20, 2018

The Banach–Tarski paradox



According to Wikipedia:
"Given a solid ball in 3‑dimensional space, there exists a decomposition of the ball into a finite number of disjoint subsets, which can then be put back together in a different way to yield two identical copies of the original ball. Indeed, the reassembly process involves only moving the pieces around and rotating them without changing their shape.

However, the pieces themselves are not "solids" in the usual sense, but infinite scatterings of points. The reconstruction can work with as few as five pieces.

A stronger form of the theorem implies that given any two "reasonable" solid objects (such as a small ball and a huge ball), the cut pieces of either one can be reassembled into the other. This is often stated informally as "a pea can be chopped up and reassembled into the Sun" and called the "pea and the Sun paradox".
I remember reading Richard Feynman's reaction to this in his memoir "Surely You're Joking Mr Feynman!",
"Then I [Feynman] got an idea. I challenged them [the mathematicians]: “I bet there isn’t a single theorem that you can tell me – what the assumptions are and what the theorem is in terms I can understand – where I can’t tell you right away whether it’s true or false.”

It often went like this: They would explain to me, “You’ve got an orange, OK? Now you cut the orange into a finite number of pieces, put it back together, and it’s as big as the sun. True or false?”

“No holes.”

“Impossible!

“Ha! Everybody gather around! It’s So-and-so’s theorem of immeasurable measure!”

Just when they think they’ve got me, I remind them, “But you said an orange! You can’t cut the orange peel any thinner than the atoms.”

“But we have the condition of continuity: We can keep on cutting!”

“No, you said an orange, so I assumed that you meant a real orange.”

So I always won. If I guessed it right, great. If I guessed it wrong, there was always something I could find in their simplification that they left out."

There are no infinities in reality.

The LTOV including rent answers critics



The Labour Theory of Value (LTOV) has not had a good press from neoclassical economists.
"However, Ricardo was troubled with some deviations in prices from proportionality with the labor required to produce them. For example, he said "I cannot get over the difficulty of the wine, which is kept in the cellar for three or four years [i.e., while constantly increasing in exchange value], or that of the oak tree, which perhaps originally had not 2 shillings expended on it in the way of labour, and yet comes to be worth £100."   [Wikipedia].
As we shall see, the solution to this mystery is the combination of socially-necessary labour time and rent.

---

Q. What is value theory in Marxism?

A. Marx distinguished three kinds of value. Firstly value itself, the amount of (socially necessary) labour time which went into the production of an object; secondly the exchange value of the produced object when it was exchanged in a market transaction for another object (which could be the universal commodity money); thirdly and qualitatively, the use value which as its name implies was the utility of the object to a person.

Q. Tell me more about value per se

A. Take a wooden box made by a carpenter by hand. Plainly a lazy, incompetent carpenter would take longer to make the box but that extra labour time could not increase its value. A given society knitted together by market transactions would soon arrive at the notion of the average period of time (hours, say) which a box like that would take to produce. That time is the socially necessary labour time. If productivity improved, the value would go down, not up.

Q. And exchange value?

A. An object which is exchanged in a market transaction is called a commodity. A commodity will typically exchange in barter for another object which took something like an equivalent time to produce (otherwise someone is getting a free ride .. and eventually more people will move into producing that commodity and compete). In any kind of competitive and established market a commodity typically exchanges for the money commodity, which makes transactions so much more fluid. The exchange value of a commodity is then its equivalent in money, which in a very simple model would be almost the same concept as its price.

The idea is that in a competitive market of freely working producers the market price will tend to oscillate - due to vagaries of supply and demand in the first instance - around the value.

It's a first, simple model.

Q. And use value?

A. This is not measured in hours or currency units. It's not a quantitative attribute. It instead registers the qualitative fact that anything socially produced has to have utility for someone, otherwise why did the producer bother. The non-trivial idea is that objects can be produced, say for the use value  of household consumption, which have a value (because they took a certain amount of time to do or make) but don't have an exchange value because they were never traded.

Q. So here's a scenario. A tropical island. The guys are doing hunting and fishing and some hard-scrabble cutivating so there's an economy and a currency of sorts. 

And then this big guy monopolises the one set of banana trees and extorts payment for bananas. How does that work in terms of value?

A. Good example. Bananas have a use value, obviously, otherwise everyone would just ignore the big guy. By hypothesis, bananas make themselves. No-one has to labour (by assumption) so the value of a banana is zero. Likewise the exchange value of a banana is also zero - as no labour is incorporated in it.

Q. But the big guy is able to charge. He makes money doesn't he?

A. So this is a case where price and exchange value differ. The big guy is extracting a rent due to his monopolising a scarce resource. We see the same with landowners, who also charge rent to access land which they have perhaps done nothing at all to improve.

If the big guy wasn't there, the bananas would be a 'windfall' and it would be first-come first-served, absent a state structure to ration and allocate (like the land runs in the nineteenth century USA).

Note that value theory can't predict how big the rent (ie the market price) will be. Who knows how much people will value a banana over other traded objects. How much disposable income they will have to allocate to bananas.

Here's the important bit. Some commodities mix aspects of value creation and rent, often when a natural process is involved such as in forestry or wine maturing. The ability of the landowner or wine-make to monopolise the resource while nature does its work allows the extraction of a rent.

This addresses Ricardo's difficulty mentioned at the top of this piece.

Q. Rent vs exchange value - a distinction without a difference?

A. No. Exchange value results from value creation; rents are simply value appropriation.

Suppose we had a guy who guards bananas (which grow themselves), a guy who guards oranges (which grow themselves), a guy who guards a bay with super-abundant fish stocks (which collect and beach themselves). Let's assume they barter. What are the right ratios for exchange?

Marx would say that all three products are not commodities, they have no value. The prices are rents - set by force majeure and need.

Assuming all three parties require all three products then if each is in local abundance the price will be zero. Each is wasting their time trying to monopolise their resource. It's like the air guys. Chill!

If the products are not in abundance only force will decide. The most aggressive guy will set whatever barter ratio he can enforce, then the next most aggressive. However, the underserved guys will eventually die of starvation, making this self-defeating. Hoarding and rationing may postpone the fateful day.

In microeconomic terms, the supply cannot be changed since it is constrained by nature (it makes itself) so the supply curve is vertical. In the absence of substitutes the demand curve is vertical too: to survive you do need a certain level of resources.

If the two vertical lines coincide we have an equilibrium (of sufficiency). If the supply line is to the right  everyone lives and some produce will rot. If it's to the left the guys will become increasingly malnourished and eventually will die.

Take a look at the diagram at the top of the post.

This simple model is also relevant to the owners of entirely robotic factories under total automation, but that's for a future post..

Sunday, November 18, 2018

The cheesecake paradox

This afternoon Clare told me: "Tonight, you'll either have a cheesecake for dessert, or you won't. It'll be a surprise."

OK, she was toying with me. She knows I am torn between desire and calories. I mentally charted my options like this.

Figure 1

But then she had second thoughts. "No," she said. "I do want you to have the cheesecake, but I also want it to be a real surprise. You'll definitely get the cheesecake tonight or tomorrow night. One of the two. But I won't tell you which."

OK, I was disappointed, I have to admit. I hate even the possibility of  delayed gratification. A new diagram took shape in my mind.


Figure 2

I'd either get the cheesecake tonight or it would be the empty plate, in which case at least I'd be getting cheesecake tomorrow.

But wait. It's meant to be a surprise. The surprise is all in the branching possibilities. If I didn't get cheesecake tonight, then I'd be in the situation below (yellow oval).

Figure 3

No surprise there! There's no branching in the yellow oval. So there's no surprise tomorrow. And that must mean I'm getting the cheesecake tonight! Excellent!


Figure 4

Oh, wait again. No branching here. So my very certainty has removed the element of surprise. So it looks like Clare's proposal to give me a surprise cheesecake has collapsed. There's no way she can do it.

No cheesecake at all. Cruel.

Figure 5

But now I'm sure I'll get nothing, I can't rule out the possibility that nevertheless she may give me a cheesecake tonight after all.

Figure 6 = figure 2

And I will be very surprised!

---

I wrote about this paradox back in 2009. I noted that despite its simplicity, according to Wikipedia, no-one has a good solution. It looks to me that the self-reference feedback is creating alternate, flip-flopping states of certainty and uncertainty.

Another example of how concepts from AI such as game trees, the perspective of an agent with cognitive states, and Kripke models of doxastic logic can illuminate problems which seem intractable from a purely logicist or philosophical standpoint.

Friday, November 16, 2018

Mathematical determinism: 2 + 2 = 4

This mini-rant on reading how Robert Plomin's book has been traduced by the usual suspects: "Genetic determinism rides again".

---
  • "So 2 + 2 = 10 (base four), 2 + 2 = 1 (mod 3). So much for determinism."

  • "Does '2 + 2' look anything like '4' to you?" (Rolls eyes).

  • "Western imperialist bias: ٢ + ۲ = ٤."

  • "If you had two apples and two oranges that's not four of anything."

  • "So two monogamous couples make a foursome? In your dreams, chauvinist!"

  • "Like white middle-aged males reduce everything to abstractions!"

  • "The voice of privilege. This is so not true for disadvantaged minorities."

  • "It depends on what '2', '+', '=' and '4' are taken to mean. They're just conventions reflecting the patriarchal power structure."

  • "An arrogant assertion presented without a shred of proof."

  • "Life is more complex and interesting than sterile theories. Get over it."

  • "Nazis/Communists/Zionists/Russians believe that."

  • "You're just saying that to make me feel uncomfortable. You're so not in touch with your own feelings."

  • "We shouldn't discriminate against people who have a problem with what you're saying."

etc etc.

---

Say something true at the margins of the Overton window and be pecked to death by spurious attacks until everyone has forgotten the original proposition but is left believing it has been convincingly refuted.

---

It's an interesting exercise, a homework assignment, to identify the fallacies in each of the 'rebuttals' above. The required concepts can be quite sophisticated: initial algebra; the formalisation of syntax and semantics in mathematical logic - wffs, valuation functions and models; the formal notions of proof and interpretation in mathematics and science.

Add a dash of sophistry and emotionalism to complete.

Wednesday, November 14, 2018

"Blueprint: How DNA Makes Us Who We Are" - Robert Plomin

Amazon link

Robert Plomin is one of the good guys, and this book is apparently a summing up, for a general audience, of his life work. It has just arrived and is on the stack.

Meanwhile Dr James Thompson of UCL has an excellent and detailed chapter-by-chapter review, from which this short excerpt.
"Chapter 11 is about the development of genome-wide association studies. Chapter 12 is a very substantial one about genetic prediction. Chapter 12 is Plomin’s real coming out: he reveals his polygenic scores for all to see. Naturally, this is a teaching opportunity, explaining the insights and the limitations of such measures. Figs 6 and 7 are worth showing again and again, if only to explain polygenic scores and their overlaps, and the fact that they provide probabilistic estimates, not certainties."
Greg Cochran has written a review at Quillette, "Forget Nature Versus Nurture. Nature Has Won".

And Toby Young, also at Quillette: "Is Sociogenomics Racist?".  No, he argues.

I'm looking forward to reading and learning. Update: here's my review.

Monday, November 12, 2018

A feminist on sex machines: Kate Devlin

Amazon link

Dr Kate Devlin's Wikipedia entry.

---

The concept of a 'machine for sex’ is a wondrous, powerful one. As the author notes, its arc aligns with all of human history. Devlin describes ‘sex toys’ from prehistory and antiquity right through to the present day.

The term ‘sex robot’ itself is rather reified. Better is the idea of an artefact with which humans may engage sexually. Such an artefact might be genetically engineered, machine fabricated or some hybrid in between. Sex devices interestingly inhabit the intersection of science and technology, genetics and evolutionary biology, and the myriad social sciences.

If a sex robot is defined as a thing, as human property, then they existed once, currently don't really exist and may exist again in the middling future. In antiquity slaves were property and presumed to have no agency. It was routine and uncontroversial for elite slave-owning males to buy and use nubile male and female slaves for sex. As Kyle Harper recounts in his “Slavery in the Late Roman World, AD 275–425”, consequences included masters becoming (stupidly) besotted with their slave sex-objects, wifely jealousy .. and unintended offspring who would join the next generation of slaves.

Antique ‘sex robots’ were, from a functional point of view, poorly implemented. Their inner drives and motivations were not aligned with their 'function’ which made ownership fraught and only manageable through sustained terror.

Devlin’s first degree was in archaeology so she will know all this. It would have been good to read about the social, ethical and even practical implications of the widespread availability of high-functioning ‘sex robots’ in antiquity. Devlin's discussion is however (pp.116-117) superficial, flagging only the usual oppressively gendered roles found in all premodern societies stabilised by male violence. She also notes a pre-Christian sexual disinhibition of which she approves.

“Turned On” is not a book of science with some feminist advocacy. It is instead a feminist tract anchored around the topic of sex-with-artefacts (p.213). What’s this for example - a mocking rebuke to the transgression of equal outcomes?
“Why do we experience things? How do the mechanisms of our bodies and brains give rise to conscious sensations? Where does that consciousness come from? You don’t have to have an answer in the Great Zombie Debate. The philosophers can’t agree on it either, Which is why you have a bunch of very clever middle-aged white men amusingly inventing words like‘ ‘zoombie’ and ‘zimboe’ to put forward their own variations on the theory. “ (p. 102)
My emphasis. It's certainly of the moment, but it jars.

Devlin tells us she is a feminist, an ideology developed to further the interests of liberal professional women like Dr Devlin. Like all ideologies, it’s protean and eclectic, cherry picking arguments which support the cause. A characteristic of ideologies is that they narrate a world their proponents wish they were in but which rarely coincides with actuality or realistic possible outcomes.

Consider objectification, something she takes strong issue with. The discussion is phenomenological: an example might be a builder wolf-whistling an attractive woman in the street; or a bunch of women hooting and laughing at a male stripper at a ‘hen party’.

In the broad-brush triune brain model, primary drives such as lust are associated with the (animalistic) brainstem formation; emotional attachment with the mammalian limbic system; and rational interaction with hominid cortical systems. It’s not much of a surprise that humans can exhibit sexual behaviour dominated by any of these loci.

Objectification (the mode of lust-dominance) would then be a brain-stem determined form of behaviour. In more refined circles we expect cortical inhibition/mediation of primary drives to deliver more measured, prosocial behaviour factoring in social context. Our biology is complicated in social settings.

There isn’t any such framing in Devlin's book. Just normative presumption that objectification is out there (for some reason), that’s it’s wrong and that current sex doll designs play up to it. This is to sell the reader short, replacing analysis with moralising.

Naturally we don't wish to succumb to the naturalistic fallacy; humans with their small group evolutionary history are imperfectly adapted to large scale societies. There are plenty of natural urges we need to regulate and indeed legislate against. There are few easy answers here as Devlin would be the first to argue, in contexts such as the legality and ethics of child sex dolls.

Devlin adopts the fashionable feminist view that phenomenal gender differences are purely social constructs. This despite the enormous weight of hard evidence (evolutionary, neuroanatomic, genomic, physical, psychometric) for well-defined and reproducible biological differences between the sexes. Differences which are hardly obscure, but recognition of which might undermine the claims of her interest group. Public choice theory assumes its usual relevance here.

Devlin interviews the CEO of RealDoll, Matt McMullen, and gives him a hard time about the overwhelming preponderance of hyper-sexualised female dolls and robots. 'Where are the less-sexualised dolls, the male dolls?' she wants to know. Actually this is a point she takes up with all the doll manufacturers she meets. The replies she gets are defensive, framed in terms of male-female differences in sexuality leading to skews in demand. Devlin is having none of it, blaming biased marketing and uncritical social conditioning (pp. 153-154). Time for a quick review of microeconomics (supply-demand equilibria?) then as I reflect on the democracy of markets in probing the world as it actually is.

Then we read this in a meandering discussion of rape fantasies and the claimed lack of any genetic influence:
“If anything, rape would theoretically reduce the reproductive success of our ancestors as it takes away selective genetic choice. “ (p. 233).
Reality is more nuanced. Rape is historically (and currently) commonplace in intergroup conflicts. It's plainly adaptive for males in the absence of draconian ingroup penalties. I suggest a quick read of Dawkins' "The Selfish Gene".

Gendered social roles have historically been oppressive to women. There are biological reasons (eg male physical strength, aggression and paternity-uncertainty) which interlink with social reasons (eg within-family inheritance) which are specific to different kinds of society and which need to be explicitly teased out. If capitalism appears to be intrinsically gender-blind in its desire to free everyone up to maximally work, then the deleterious effects on human self-reproduction also need some analysis.

A perennial science-fictional trope referenced by Tom Whipple, The Times science editor in his review, is that of the perfected sex robot as a kind of sterile mosquito (which has already resulted in some local extinctions of this malaria disease vector).

This is not a problem we'll face anytime soon but is there something to it? Devlin is unworried while Whipple remains concerned. Plainly it’s hard to assess an unknown artefact but with universal and easy access to contraception in the west, perhaps we already have a natural experiment. Check those Total Fertility Ratios.

We are already selecting for women who positively want children (rather than just sex); ubiquitous effective and sterile sex robots would select for men with a similar drive. Let's hope there's that much variability in the gene pool.

In summary Devlin's book is an easy and amusing read: somewhat superficial; an interesting tour of the sex doll/sex toy landscape which will be unfamiliar to many readers; intriguing confessional snippets from the author's private life.

The book is not particularly scandalous or salacious, it's not very conceptual or analytic and its opinions are conventional liberal left. Put aside a slightly sprawling and uneven structure and it reads like an extended New Scientist article.

Incidentally, if Dr Devlin or her co-workers were ever to read this review, they would not be won over. It is in the nature of ideologies to be believed by their adherents and they provide powerful mechanisms of contextualisation and framing to theorise their opponents. This review would be framed as a defence of the patriarchal power structure: discredited gender essentialism.

We are truly the victims of our own axioms.

---

Further reading: "So Beautiful".

Sunday, November 11, 2018

A second review of "The Master Algorithm" - Pedro Domingos

Amazon link

This is a guest review from Dr Roy Simpson.

---

Review of The Master Algorithm (by Pedro Domingos) 

By Dr. R. Simpson

This book provides both a history and a visionary project in machine learning by a leading professor in the field. The author writes fluently and well providing an informative book, which can be worth re-reading if one is interested in the details of machine learning techniques, as well as in (re-)evaluating his ideas about the future of the subject.

I was brought to this book after reading two other recent books on the subject of Algorithms: Weapons of Math Destruction, by Cathy O'Neil and the recent Hello World by Hannah Fry. Both of these books discuss the ongoing issues associated with the application of Algorithms, Big Data and Machine Learning to contemporary society – the latter book being newer and most relevant for UK audiences. These books recommended The Master Algorithm (2015) as a book for a deeper understanding of machine learning.

Much of this book is eminently quotable, and its text provides good introductions, for example here is from the first page of the Prologue:

You may not know it, but machine learning is all around you. When you type a query into a search engine, it's how the engine figures out which results to show you (and which ads as well). When you read your email, you don't see most of the spam, because machine learning filtered it out. Go to Amazon.com to buy a book or Netflix to watch a video, and a machine-learning system helpfully recommends some you might like...

Traditionally, the only way to get a computer to do something – from adding two numbers to flying an airplane - was to write down and algorithm explaining how, in painstaking detail. But machinelearning algorithms, also known as learners, are different: they figure it out on their own, by making inferences from data. And the more data they have, the better they get. Now we don't have to program computers; they program themselves.

It's not just in cyberspace, either: your whole day, from the moment you wake up to the moment you fall asleep, is suffused with machine learning.

The introductory chapters continue with more aspects of machine learning making some interesting points. For example the above passage indicates a “culture shift” within the programming world. In a later section it is noted that Microsoft has some difficulties with the new world, because its programmers are just that and its main products are produced in the traditional way, whereas Google is more of a machine learning organisation, with its main products produced the machine learning way.

He introduces the metaphor that whereas traditional programming is more like a manual (at best industrial) process, machine learning is more like farming – prepare the ground, then sit back and watch the systems (i.e. commercial products) grow themselves.

The author is aware that this field is not “General AI” - a topic often discussed in this blog to which I shall return at the end of this review. However we read that within this subfield of AI the traditional opponents are “Knowledge Engineers”. Knowledge Engineers hold (or at least once held) the view that systems which contain “knowledge” need to have that knowledge typed into them, and more generally be programmed in the traditional way.

The classic example was Cyc - a massive common sense knowledge based system over which programmers have been typing in “common sense rules” for several decades now.

The main counter to this view presented in this book is the matter of scale: knowledge engineers could work with thousands of rules; whereas machine learning will generate many millions of rules. (The argument against this knowledge engineering viewpoint held by the Agent Theory/AGI community has been different, and concerns its narrowness of scope. Agent Theorists might have used a “narrowness of scope” argument against early machine learning too, but the scalability of these techniques has won the commercial argument  - at least for now - though not necessarily the conceptual argument, as discussed later in this review.)

The primary content of the book is a review and overall analysis, from a modern machine learning perspective, of the five main strands of machine learning during the history of AI: Symbolism, Connectionism, Evolutionary Programming, Bayesian Networks, and Analogical Reasoning.

His purpose in all of this is to identify the principles involved and describe the essence of the technique and identify the best algorithm that the given technique has provided to the machine learning community. From there he moves to the main objective of the book: the development of a “Master (Learning) Algorithm” which incorporates all the best of the previous techniques and which is the subject of his own research group. This Master Algorithm would then be able to optimally learn anything .. .

A brief summary of each (with quotes from the book):

Symbolism

The symbolist's core belief is that all intelligence can be reduced to manipulating symbols. Whereas deductive reasoning is about going from axioms to conclusions, the learning aspect requires inductive reasoning: going from conclusions to axioms. Thus Hume's problem is discussed and the master algorithm here becomes “inverse deduction”.

Inverse deduction is like a super-scientist systematically looking at the evidence, considering possible inductions, collating the strongest, and using those along with other evidence to construct yet further hypotheses – all at the speed of computers. Yet this inverse deduction has limitations and issues, so we move to Connectionism.

Connectionism

How does your brain learn? Hebb's rule (from 1949) has become the cornerstone of connectionism, and is about neuron firing: “neurons that fire together wire together”. This history in AI begins with the Perceptron (from the 1950s). This was an electromechanical learning machine, based on a model of neurons, which had some basic classification skills (e.g is this a picture of a door or not?).

In the late 1960s it was proven that its classification skills were too limited to be of much generality and this approach to AI learning suffered a near total set-back (to the delight of the knowledge engineers of the time, due to funding competition).

It was not until the 1980s that physics inspired alternatives were introduced, such as the Boltzmann machine, which introduced probabilities into the field, and brought in techniques from statistical physics (ie thermodynamics) – so for a period the notion of “temperature” was important for a neural network learning system!

Although the introduction of probabilities has lasted, it is not clear what has happened to "temperature" in machine learning. It later transpired also that none of this physics theory was necessary to overcome the original Perceptron limitations. Nevertheless more mathematical techniques emerge from this era such as calculations in hyperspace, with associated weighted functions, convergence metrics, etc.

The main algorithm eventually identified from Connectionism is backpropagation, whose refinements are at the core of today's learning systems. However, the author raises an intriguing question: Is everything we “know” actually learned by our neurons? Has not evolution played a part too?

Evolutionary Learning

As an introduction to this chapter the author tells the following fantasy story:

Robotic Park is a massive robot factory surrounded by ten thousand square miles of jungle, urban and otherwise. Ringing that jungle is the tallest, thickest wall ever built, bristling with sentry posts, searchlights, and gun turrets. The wall has two purposes: to keep trespassers out and the park's inhabitants – millions of robots battling for survival and control of the factory – within.

The winning robots get to spawn, their reproduction accomplished by programming the banks of 3D printers inside. Step-by-step, the robots become smarter, faster – and deadlier. Robotic Park is run by the US Army, and its purpose is to evolve the ultimate soldier.

Needless to say this story brings out a large number of issues and concerns. The author seems to be sanguine about the dangers here, although modern “AI Ethics” movements (which included the late Stephen Hawking and Elon Musk) are concerned about this type of development. Towards the end of the book, the author actually suggests that working on “AI Ethics” could itself be a growth industry for displaced humans in the new era – robots and AI will have a lot of (human) ethics to learn!

The author is using this story here to dramatically introduce the AI Learning techniques inspired by Darwinian evolution: genetic algorithms and genetic programming. These techniques are overviewed, and some issues identified. However this reviewer is intrigued by the wider point that is being made here, and has an alternative way of expressing a related idea. The overall point that the author is making can be summarised by this formula:

   Learning == Agent Learning + Environment Learning + (Environment → Agent transfer)

In other words, when we ask “how does that small brain learn all this stuff?” - the answer is that the small brain has not had to do all the learning implicit in its actions. This viewpoint is implied also by the Chomsky view of language acquisition (Chomsky has been another critic of the traditional approach to machine learning – the author hopes that his wider “Master Algorithm” approach meets Chomsky's concerns.)

There are again several limitations to any master algorithm provided by what we could also call the Darwinian algorithm. Chief amongst them is the fact that the Darwinian algorithm (as currently understood) tends to find “suboptimal” solutions, not optimal solutions – so we move on to Bayesian theory.

Bayesian Networks

The path to optimal learning begins with a formula that many people have heard of: Bayes Theorem. But here we'll see it in a whole new light and realize that it's vastly more powerful than you'd guess from its everyday uses.

At heart, Bayes' theorem is just a simple rule for updating your degree of belief in a hypothesis when you receive new evidence: if the evidence is consistent with the hypothesis, the probability of the hypothesis goes up; if not, it goes down.

Bayes Theorem uses conditional probabilities. (P(A|B) is the probability that A happens given that B has happened/is assumed) and there is an associated inference system. Similar to deductive logic (e.g. the Prolog-based resolution systems sometimes discussed in this blog) the inference system allows the deduction of probabilistic conclusions and the management of probabilistic assumptions. This is all wrapped into a network structure amongst assumptions, and apparently it has been discovered that there is an isomorphism between the probabilities and weights often used in Artificial Neural Net models and the probabilities used in Bayesian Networks.

The success of this Bayesian approach arises directly from the quantity of data now available, making the probabilities and conditional probabilities very accurately determinable from millions upon millions of data points (pre-Internet such probabilities would have been input by researchers by hand as guesses, resulting in unconvincing performance and results from such probabilistic systems).

This has motivated much of the idea behind a complete Master Algorithm for learning. This has left open two remaining unification tasks, of which the hardest has been the unification of logic and probability; the other has been to incorporate what to do when there is essentially no data, only analogies to existing data.

Analogy Reasoning

Methane and Methanol have very similar chemical structure; but they are not identical and have some big differences since one is a (room temperature) gas, the other a liquid. So if your system knew much about one, what can it deduce about the other? Much of science and business-services-work operates by analogy: no two customers are exactly the same. We manage to cope, so what are the principles involved?

Various classification ideas have been developed in this strand of AI, and a very powerful technique called a Support Vector Machine was apparently the most powerful AI technique around the turn of the century. Only in recent years has it been superseded by Artificial Neural Nets for the top learning slot (although Artificial Neural Nets don't do analogical reasoning as such).

The Master Algorithm - Alchemy

So putting this all together has been the research project of Prof. Domingos, and he has developed a mathematical framework called Markov Logic Networks and a corresponding (open source) software system called Alchemy. The final chapters discuss the general steps required to produce a Master Algorithm from all the above ingredients. He does not describe the current form of Alchemy (or MLN) as the Master Algorithm, but uses it to suggest that this goal is feasible.

One feature he seems to claim Alchemy is lacking is the ability to explain itself fully, also there may still be optimisation issues. There are also some discussions about the social aspects of all this, suggesting that users should form “trusted data unions” to hold the value of their data (a bit like banks holding their money), rather than the current practice of just handing everything over to the big corporations who then benefit from any commercial consequences (and as UK residents are aware, don't pay appropriate tax either).

Analysis: Agent Theory and Computational Mathematics

I shall end this review with a few remarks about this idea of a “Master Learning Algorithm” from the perspective of Agent theory (which is often discussed in this blog); and some ideas about Computational Mathematics (my own interest).

Agent theory takes a more holistic, biologically realistic and physically realist view of AI than one finds in stand-alone techniques like machine learning. Although these machine learning techniques have become (and will remain for another decade at least) at the centre of a business and social revolution, in aiming for such a high goal as a Master Algorithm which can learn everything – and do so optimally - one has to ask if boundaries will eventually be reached due to insufficient attention to agent theoretic issues.

For example, how and in what way, is a biological cell a learning system? Likewise the mysteries of neurological biochemistry are not resolved in an agreed way. Even aspects of the Darwinian algorithm are not fully understood. So it is quite possible that nature has more to teach us about learning.

On the subject of Computational Mathematics and AI much can also be said. A famous critic of the entire AI project (at least as seen as a branch of Computer Science) has been Professor Penrose with his books (from 1989 and 1994). Since that time more results have appeared in the foundations of logic and mathematics which can be viewed as clarifying Penrose's position.

For example using the results of an ongoing project known as ReverseMathematics, this reviewer suspects that much of Applied Engineering Mathematics is not actually computable in the Turing sense. However it nearly is, so many effects are “subtle” - although the inability to predict weather systems accurately beyond (say) five days; turbulence theory; and mathematical puzzles related to physics models suggest these “weakly non-computable” effects may yet be important.

The author himself recognises that an outstanding and unsolved mathematical problem known as the NP-completeness problem is significant to the subject of AI. I shall leave the last words to the author:

The purpose of AI systems is to solve NP-complete problems, which may take exponential time, but the solutions can always be checked efficiently. We should therefore welcome with open arms computers that are vastly more powerful than our brains, safe in the knowledge that our job is exponentially easier than theirs.

---

My own review was posted back in June 2016 and can be found here.