Showing posts with label embodied robotics. Show all posts
Showing posts with label embodied robotics. Show all posts

Saturday, April 01, 2017

Google Translate: English to predicate logic (please!)



Did you see the  "Missing: google"?

---

A big problem with English (natural language really) is that it doesn't come equipped with an explicit set of inference rules. Consequently, when someone uses natural language to communicate with an AI system, it's not really possible for that system to immediately connect the utterance to its store of knowledge. If only natural languages were like formal languages, which have proper inference and well-defined semantics. The thought that secretly they are was the intuition of Richard Montague*. But he was misguided.

Any AI natural language understanding system tries to transform the raw material of human language into something it can use, something more inferentially tractable.  Usually that doesn't work too well, and even the latest statistical systems (which do well in surface-level speech-recognition and translation) show scant abilities to understand.

It's as well to remind ourselves just why natural languages are so unhelpful to AI designers. It's because they are a highly-optimised solution to a situated communications problem. Speech is a low-bandwidth, linear and slow channel for communicating time-critical thoughts. So speech is highly optimized to use every available constraint to speed up meaning transfer:

  • volume, pitch, timbre and tone of voice
  • shared and predictive knowledge of the conversational partner
  • emotional cues
  • physical gesturing and facial expressions
  • environmental situation and context 
  • ...


Researchers are quite aware of this, of course. The topic area is called Pragmatics and it's hived off as a separate sub-discipline .. because it seems to require way too much modelling of the conversing agents in their specific environment, culture and history. In short, it's too hard.

But by abstracting away these additional constraints which channel and constrain meaning, we make the semantic understanding problem way too hard. Which is why we can't solve it.

Google Translate system which mapped between a natural language and a formal language (with well-defined inference rules and semantics) would nevertheless be a boon to the designers of conversational AI systems, including chatbots. But Google doesn't have a corpus of First-Order Predicate Calculus sentences translationally-linked to English, so its deep learning systems can't crunch the data and add FOPC to its list of languages. Projects such as Cyc have attempted to do this stuff by hand .. with surprisingly little impact.

Again the way forward is embodied robotics and human baby conversational emulation.

---

* In a weird reprise of Alan Turing's fate, Wikipedia reports that Richard Montague 'died violently in his own home; the crime is unsolved to this day. Anita Feferman and Solomon Feferman argue that he usually went to bars "cruising" and bringing people home with him. On the day that he was murdered, he brought home several people "for some kind of soirée", but they instead robbed his house and strangled him.'

He was 40.

Friday, March 31, 2017

Naive generate-and-test won't hack it

When I was young I toyed with the following idea.

Pretty much any concept can be adequately expressed in a mini-essay of a thousand words.

Simply generate all possible articles of a thousand words and somewhere you will find the answer to all problems.

Want the design of a stardrive engine? Immortality? The theory of perfect governance?

It's all in there somewhere.

---

How many essays though? Apparently the average educated speaker of English knows about 40,000 words. So for our first estimate, we could simply raise 40,000 to the power of 1,000 .. but most of those 104,602 essays would be wildly ungrammatical. We can do better.

I reviewed a sample text: the introductory quote in Peter Seibel's "Practical Common Lisp".



The first five sentences comprised 100 words in total which broke down into:
  • nouns: 20%
  • verbs: 15%
  • adjectives: 10%
  • others: 55%
A certain amount of hand-wavy rounding of course. Assume we adopt the very restrictive constraint of exactly one syntactic structure for the entire set of essays, then the total number reduces to a product of:
(number-of-English-words-in-category) (number-of-words-of-this-category-in-essay)
or,
8,000200 * 6,000150 * 4,000100 * 22,000550 = 104,092
That's still a big number*. Suppose only one 'essay' in a billion was semantically sensible and we could read one essay per second. That's 104,083 seconds .. or 3 * 104,066 billion years.

The merits of a compact notation.

---

Exhaustive search through the space of all possible candidates isn't a very good way of proceeding. And this has important implications for DARPA's third wave - contextual AI - which I wrote about previously.

In his excellent exposition (YouTube), John Launchbury highlighted the very large number of training instances needed to force convergence for today's artificial neural networks. By comparison, children learn new concepts from very few examples.

John Launchbury's proposed solution was - correctly - to identify additional constraints which might dramatically collapse the search space. His chosen example showed the benefits of adding the dynamics of handwriting characters to the resultant bitmaps normally used for training. It turns out that if you consider how the image might have been created, it makes recognition a lot easier.

It's not hard to identify the extra constraints about the world which children use. They interact with new objects, touch them, throw them, bite them and try to break them. Thus are acquired notions of 3D structure, composition and texture to augment what their visual systems are telling them.

I really do think that a high priority should be given to embodied robotics in the next wave of AI research.

---

Another example John Launchbury discussed was the Microsoft Internet-chatbot "Tay".



Apparently this was the least-offensive tweet Launchbury could find. But what would an AI have to know about contemporary mores to self-reject statements like that?

For extra credit, discuss the 'situated cognition' thesis that only through active and corporeal participation in the social world can one truly understand social concepts.

Particularly emotionally-charged ones.

---

* Since
(i)  I don't consider all the syntactically-permissible permutations of the ways in which nouns, adjectives, verbs and others could be mixed up in the thousand words, while

(ii)  the size of the 'others' vocabulary is likely to be way smaller than 22,000 (so if, for example, the 'others' vocabulary size was 2,200, this would reduce the overall essay-set size by a factor of 10550 - a distinction, however, without a practical difference),
this calculation counts as pretty bogus. I only wanted to demonstrate, however, that no matter how you cut it, the numbers involved are simply ginormous.