Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Just because you chose Markov chains as your modeling mechanism doesn't mean that there is no statistical modeling method that is capable of developing something passing for what we'd call "meaning".

This is the same argument that was used against artificial neural networks. Neural network of type A can't do X, therefore neural networks will never do Y.

Language is immensely complex, and real human language involves things which are not encoded in text (and i'd remind you that you were trying to infer meaning from text specifically, not the full multi-channel robustness of humans communicating), we don't even have a full handle on what all of the cognitive processes and factors are that go into the production and understanding of language (although we've developed a lot of interesting work to those ends).

So hearing folks give up claim that Chomsky is correct because our current tools aren't up to the job is a bit puzzling, because we don't even have a complete understanding of what sort of thing language is or what sorts of things we are as systems which can use language.

Chomsky has opinions (and some facts) about what language is, and we are, but he does not have solid proof to confirm his specific conjectures. Is human language context free? context sensitive? Something else? (Chomsky's minimalist program uses movement along a tree to preserve referentiality and a bunch of junk, alternative syntactic frameworks such as HPSG uses directed graphs as the basis of their language modeling. Still others do weirder things like higher order combinatoric logics. And unfortunately none of the theoretical frameworks appear to be without their drawbacks)



I am not a specialist, but as far as I know, Chomsky's argument here was that the existence of recursion showed that a Markov approach had to be wrong. Surely a similar argument can be made for statistical approaches? There is no way to represent a reference to some other part of the statement in a purely statistical method. If they work they happen to work basically by accident.

Just blue-skying here, but it seems to me that if I knew enough about how a statistical program worked, I could craft a sentence that would utterly confuse it, even though it was perfectly intelligible to a normal English speaker. A putative strong-AI program could not be fooled in this way.


Except that his argument is somewhat moot as a practical matter, because there are no infinitely recursive sentences (given that all sentences are finite).

Long distance dependencies are an issue in language modeling that do need to be accounted for, but all that tells me is that Markov chains aren't the right structure to model language (unless, maybe you had a MASSIVE amount of data, and a markov chain of an order high enough that you account for the majority of sentences. maybe).


You can statistically build a model that has recursion. It's just that such a model cannot be sure it has induced the right grammar - that's what Chomsky's argument was. I think the obvious counter argument is so what? Given any other constraints like parsimony you can certainly reliably induce a grammar.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: