Das Modul scikits. timeseries ist nicht mehr aktiv. Es gibt eine hervorragende Liste von Bugs, die wahrscheinlich nicht behoben werden. Der Plan ist, dass die Kernfunktionalität dieses Moduls in Pandas umgesetzt wird. Wenn du dieses Modul sehen möchtest, das unabhängig von Pandas lebt, fühlst du dich frei, den Code zu geben und ihn zu übernehmen. Das Modul scikits. timeseries bietet Klassen und Funktionen zum Manipulieren, Melden und Plotten von Zeitreihen verschiedener Frequenzen. Der Fokus liegt auf dem bequemen Datenzugriff und der Manipulation, während die vorhandene mathematische Funktionalität in numpy und scipy genutzt wird. Wenn die folgenden Szenarien Ihnen vertraut sind, dann finden Sie wahrscheinlich das scikits. timeseries Modul nützlich: Vergleichen Sie viele Zeitreihen mit verschiedenen Datenbereichen (zB Aktienkurse) Erstellen Sie Zeitreihenplots mit intelligent beabstandeten Achsetiketten Konvertieren Sie eine Tageszeitreihe Bis monatlich, indem man den durchschnittlichen Wert während jedes Monats einnimmt Arbeit mit Daten, die fehlende Werte haben Bestimmen Sie den letzten Werktag des Vormonats Quartals für Berichtszwecke Berechnen Sie eine bewegte Standardabweichung effizient Dies sind nur einige der Szenarien, die mit den Scikits sehr einfach gemacht werden Zeitmodul. Dokumentation1 Sprachverarbeitung und Python Es ist leicht, unsere Hände auf Millionen von Worttexten zu bekommen. Was können wir damit machen, unter der Annahme, dass wir einige einfache Programme schreiben können In diesem Kapitel wenden wir uns bitte an folgende Fragen: Was können wir erreichen, indem wir einfache Programmiertechniken mit großen Textmengen kombinieren. Wie können wir automatisch Schlüsselwörter und Phrasen extrahieren, Stil und Inhalt eines Textes Welche Werkzeuge und Techniken macht die Python-Programmiersprache für eine solche Arbeit. Was sind einige der interessanten Herausforderungen der Verarbeitung natürlicher Sprache Dieses Kapitel ist in Abschnitte unterteilt, die zwischen zwei ganz anderen Stilen überspringen. Im Zitat mit Sprachquotten werden wir einige sprachlich motivierte Programmieraufgaben übernehmen, ohne dass wir uns zwingen müssen, wie sie funktionieren. In der näheren Betrachtung von Pythonquot-Abschnitten werden wir die Schlüsselprogrammierungskonzepte systematisch überprüfen. Nun markieren Sie die beiden Stile in den Abschnitt Titel, aber später Kapitel werden beide Stile zu mischen, ohne so up-front darüber. Wir hoffen, dass dieser Einführungsstil Ihnen einen authentischen Geschmack dessen gibt, was später kommen wird, während wir eine Reihe von Elementarkonzepten in der Linguistik und Informatik abdecken. Wenn Sie grundlegende Vertrautheit mit beiden Bereichen haben, können Sie zu 5 überspringen, werden wir alle wichtigen Punkte in späteren Kapiteln wiederholen, und wenn Sie irgendetwas vermissen, können Sie das Online-Referenzmaterial bei nltk. org leicht konsultieren. Wenn das Material völlig neu für Sie ist, wird dieses Kapitel mehr Fragen aufwerfen als es beantwortet, Fragen, die im Rest dieses Buches angesprochen werden. 1 Informatik mit Sprache: Texte und Worte waren alle sehr vertraut mit Text, da wir jeden Tag lesen und schreiben. Hier werden wir Text als Rohdaten für die von uns geschriebenen Programme behandeln, Programme, die es in einer Vielzahl von interessanten Weisen manipulieren und analysieren. Aber bevor wir das machen können, müssen wir mit dem Python-Dolmetscher beginnen. 1.1 Erste Schritte mit Python Eines der freundlichen Dinge über Python ist, dass es Ihnen erlaubt, direkt in den interaktiven Interpreter 8212 das Programm, das Ihre Python-Programme ausgeführt wird, eingeben. Sie können auf den Python-Interpreter zugreifen, indem Sie eine einfache grafische Oberfläche namens Interactive DeveLopment Environment (IDLE) verwenden. Auf einem Mac finden Sie das unter Applikationen 8594 MacPython. Und unter Windows unter Alle Programme 8594 Python. Unter Unix kannst du Python aus der Shell ausführen, indem du im Leerlauf tippst (falls dies nicht installiert ist, probierst du python aus). Der Dolmetscher druckt eine Klappe über deine Python-Version, um zu überprüfen, ob du Python 3.2 oder später betreibst (hier ist es für 3.4.2): Das erste Mal, wenn du eine Konkordanz in einem bestimmten Text verwende, dauert es noch ein paar Sekunden Ein Index, so dass nachfolgende Recherchen schnell sind. Ihre Umdrehung: Versuchen Sie, nach anderen Wörtern zu suchen, um die Wiederholung zu speichern, können Sie den Pfeil nach oben, Pfeiltaste oder Alt-p verwenden, um auf den vorherigen Befehl zuzugreifen und das gesuchte Wort zu ändern. Sie können auch bei einigen der anderen Texte suchen, die wir enthalten haben. Zum Beispiel suche Sinn und Sinn für das Wort Zuneigung. Mit text2.concordance (quotaffectionquot). Suchen Sie das Buch der Genesis, um herauszufinden, wie lange einige Leute lebten, mit text3.concordance (quotlivedquot). Du könntest Text4 anschauen. Die Eröffnungsadressenkorpus. Um Beispiele für Englisch zu sehen, die bis 1789 zurückgehen und nach Worten suchen wie Nation. Terror. Gott, um zu sehen, wie diese Wörter im Laufe der Zeit anders verwendet wurden. Weve auch enthalten text5. Der NPS Chat Corpus. Suche dies für unkonventionelle Wörter wie im. Ur Lol (Beachten Sie, dass dieses Korpus unzensiert ist) Sobald Sie ein paar Zeit damit verbracht haben, diese Texte zu untersuchen, hoffen wir, dass Sie ein neues Gefühl für den Reichtum und die Vielfalt der Sprache haben. Im nächsten Kapitel erfahren Sie, wie Sie auf eine breitere Palette von Text zugreifen können, einschließlich Text in anderen Sprachen als Englisch. Eine Konkordanz erlaubt uns, Worte im Kontext zu sehen. Zum Beispiel haben wir gesehen, dass ungeheuerlich in Kontexten wie den Bildern und einer Größe aufgetreten ist. Welche anderen Worte erscheinen in einem ähnlichen Kontextbereich. Wir können herausfinden, indem wir den Begriff ähnlich dem Namen des Textes anhängen und dann das entsprechende Wort in Klammern einfügen: 2 Ein näherer Blick auf Python: Texte als Wörter der Wörter Nach dem Prompt weve gegeben einen Namen, den wir gemacht haben, sent1. Gefolgt von dem Gleichheitszeichen, und dann einige zitierte Wörter, getrennt durch Kommas, und umgeben von Klammern. Dieses Klammermaterial ist als Liste in Python bekannt: So speichern wir einen Text. Wir können es überprüfen, indem wir den Namen eingeben. Wir können um seine Länge bitten. Wir können sogar unsere eigene lexicaldiversity () - Funktion dazu anwenden. Beachten Sie, dass unsere Indizes von Null beginnen: Sende Element Null, geschrieben sent0. Ist das erste wort, word1 Wohingegen das gesendete Element 9 Wort 10 ist. Der Grund ist einfach: Der Moment, in dem Python auf den Inhalt einer Liste aus dem Computer-Speicher zugreift, ist es bereits beim ersten Element, dass wir ihm sagen müssen, wie viele Elemente vorwärts gehen. So verlässt es null Schritte vorwärts beim ersten Element. Diese Praxis des Zählens von Null ist zunächst verwirrend, aber typisch für moderne Programmiersprachen. Youll schnell schnell den Hang davon, wenn Sie beherrschen das System der Zählung Jahrhunderte, wo 19XY ist ein Jahr im 20. Jahrhundert, oder wenn Sie leben in einem Land, wo die Böden eines Gebäudes sind von 1 nummeriert, und so zu Fuß n-1 Treppen führt Sie auf Stufe n. Nun, wenn wir versehentlich einen Index verwenden, der zu groß ist, bekommt man einen Fehler: Diesmal ist es kein Syntaxfehler, da das Programmfragment syntaktisch korrekt ist. Stattdessen ist es ein Laufzeitfehler. Und es erzeugt eine Traceback-Nachricht, die den Kontext des Fehlers zeigt, gefolgt von dem Namen des Fehlers, IndexError. Und eine kurze Erklärung. Lass uns einen genaueren Blick auf das Schneiden werfen, mit unserem künstlichen Satz wieder. Hier verifizieren wir, dass die Scheibe 5: 8 gesendete Elemente in den Indizes 5, 6 und 7 enthält: Wir können ein Element einer Liste ändern, indem wir einem der Indexwerte zuordnen. Im nächsten Beispiel setzen wir auf die linke Seite des Gleichheitszeichens. Wir können auch eine ganze Scheibe mit neuem Material ersetzen. Eine Konsequenz dieser letzten Änderung ist, dass die Liste nur vier Elemente hat und der Zugriff auf einen späteren Wert einen Fehler erzeugt. Ihr Zug: Nehmen Sie sich ein paar Minuten Zeit, um einen eigenen Satz zu definieren und einzelne Wörter und Wortgruppen (Scheiben) mit den gleichen Methoden zu ändern, die früher verwendet wurden. Überprüfen Sie Ihr Verständnis, indem Sie die Übungen auf Listen am Ende dieses Kapitels ausprobieren. 2.3 Variablen Ab Anfang 1. haben Sie Zugriff auf Texte namens text1. Text2 und so weiter. Es hat viel geschrieben, um in der Lage zu sein, auf ein 250.000-Wort-Buch mit einem kurzen Namen wie dies zu verweisen. Im Allgemeinen können wir Namen für alles, was wir kümmern, um zu berechnen. Wir haben dies in den vorigen Abschnitten, z. B. Definieren einer Variablen sent1. Wie folgt: Solche Linien haben die Form: Variable Ausdruck. Python wird den Ausdruck auswerten und sein Ergebnis auf die Variable speichern. Dieser Vorgang wird als Zuweisung bezeichnet. Es gibt keine Ausgabe, die Sie die Variable auf eine eigene Zeile eingeben müssen, um ihren Inhalt zu überprüfen. Das Gleichheitszeichen ist etwas irreführend, da sich die Information von der rechten Seite nach links bewegt. Es könnte helfen, daran zu denken, wie ein Pfeil nach links. Der Name der Variablen kann alles sein, was Sie mögen, z. B. Mysent Satz Xyzzy Es muss mit einem Brief beginnen und kann Zahlen und Unterstriche enthalten. Hier sind einige Beispiele für Variablen und Zuordnungen: Denken Sie daran, dass kapitalisierte Wörter vor Kleinbuchstaben in sortierten Listen erscheinen. Beachten Sie im vorherigen Beispiel, dass wir die Definition von mysent über zwei Zeilen aufgeteilt haben. Python-Ausdrücke können über mehrere Zeilen aufgeteilt werden, solange dies in jeder Art von Klammern geschieht. Python nutzt die quot. Um anzuzeigen, dass mehr Input erwartet wird. Es spielt keine Rolle, wie viel Einrückung in diesen Fortsetzungslinien verwendet wird, aber einige Einrückung macht sie in der Regel einfacher zu lesen. Es ist gut, sinnvolle Variablennamen zu wählen, um dich zu erinnern 8212 und jedem anderen zu helfen, der deinen Python-Code 8212 liest, was dein Code zu tun ist. Python versucht nicht, Sinn für die Namen zu geben, die es blind Ihren Anweisungen folgt und nicht widerspricht, wenn Sie etwas verwirrendes tun, wie ein zwei oder zwei 3. Die einzige Einschränkung ist, dass ein Variablenname kein Pythons reservierte Wörter sein kann, wie zB def. ob . Nicht Und importieren Wenn Sie ein reserviertes Wort verwenden, erzeugt Python einen Syntaxfehler: Achten Sie auf Ihre Wahl von Namen (oder Bezeichnern) für Python-Variablen. Zuerst sollten Sie den Namen mit einem Brief beginnen, wahlweise gefolgt von Ziffern (0 bis 9) oder Buchstaben. So ist abc23 gut, aber 23abc wird einen Syntaxfehler verursachen. Namen sind Groß-und Kleinschreibung, was bedeutet, dass myVar und myvar sind verschiedene Variablen. Variablennamen können keinen Leerraum enthalten, aber Sie können Wörter unter Verwendung eines Unterstrichs trennen, z. B. Myvar Achten Sie darauf, dass Sie keinen Bindestrich anstelle eines Unterstrichs einfügen: my-var ist falsch, da Python das Zitat als Minuszeichen interpretiert. 2.4 Strings Einige der Methoden, mit denen wir auf die Elemente einer Liste zugreifen, funktionieren auch mit einzelnen Wörtern oder Strings. Beispielsweise können wir einer Variablen einen String zuordnen. Einen String indexieren Und schneide einen String: Wir können auch Multiplikation und Addition mit Strings durchführen: Wir können den Worten einer Liste beitreten, um einen einzelnen String zu machen oder einen String in eine Liste zu teilen, wie folgt: Wir kommen zum Thema Strings in 3. Vorläufig haben wir zwei wichtige Bausteine 8212 Listen und Strings 8212 und sind bereit, auf eine Sprachanalyse zurückzukehren. 3 Computing mit Sprache: Einfache Statistik Lets zurück zu unserer Erforschung der Möglichkeiten, wie wir unsere rechnerischen Ressourcen auf große Mengen von Text zu tragen bringen können. Wir begannen diese Diskussion in 1. und sahen, wie man nach Worten im Kontext sucht, wie man das Wortschatz eines Textes kompiliert, wie man zufälligen Text im selben Stil erzeugt und so weiter. In diesem Abschnitt nehmen wir die Frage auf, was einen Text unterscheidet, und verwenden Sie automatische Methoden, um charakteristische Wörter und Ausdrücke eines Textes zu finden. Wie in 1. können Sie neue Funktionen der Python-Sprache ausprobieren, indem Sie sie in den Dolmetscher kopieren und Sie werden im folgenden Abschnitt systematisch über diese Funktionen informieren. Bevor Sie weiter fortfahren, können Sie Ihr Verständnis des letzten Abschnitts überprüfen, indem Sie die Ausgabe des folgenden Codes vorhersagen. Sie können den Dolmetscher verwenden, um zu prüfen, ob Sie es richtig gemacht haben. Wenn Sie nicht sicher sind, wie diese Aufgabe zu tun, wäre es eine gute Idee, den vorherigen Abschnitt zu überprüfen, bevor Sie weiter fortfahren. 3.1 Frequency Distributionen Wie können wir automatisch die Worte eines Textes identifizieren, die am meisten informativ über das Thema und das Genre des Textes sind. Stellen Sie sich vor, wie Sie die 50 häufigsten Wörter eines Buches finden könnten. Eine Methode wäre, eine Tally für jeden Vokabular zu halten, wie in 3.1. Die Tally würde Tausende von Reihen brauchen, und es wäre ein überaus mühsamer Prozess, der so mühsam war, dass wir die Aufgabe einer Maschine zuordnen würden. Abbildung 3.1 Zählen von Wörtern in einem Text (eine Häufigkeitsverteilung) Die Tabelle in 3.1 ist als Häufigkeitsverteilung bekannt. Und es sagt uns die Häufigkeit jedes Wortschatzes im Text. (Im Allgemeinen könnte es jede Art von beobachtbarem Ereignis zählen.) Es ist ein quittiertes Verteilungsziel, weil es uns sagt, wie die Gesamtzahl der Wortmarken im Text über die Vokabeln verteilt wird. Da wir oft Häufigkeitsverteilungen in der Sprachverarbeitung benötigen, bietet NLTK eine integrierte Unterstützung für sie. Wir können einen FreqDist benutzen, um die 50 häufigsten Worte von Moby Dick zu finden: Wenn wir zum ersten Mal FreqDist anrufen. Wir übergeben den Namen des Textes als Argument. Wir können die Gesamtzahl der Wörter (quotoutcomesquot), die bis 8212 260.819 im Fall von Moby Dick gezählt worden sind, inspizieren. Der Ausdruck mostcommon (50) gibt uns eine Liste der 50 am häufigsten vorkommenden Typen im Text. Dein Turn: Versuche das vorherige Häufigkeitsverteilungsbeispiel für dich selbst, für text2. Seien Sie vorsichtig, die richtigen Klammern und Großbuchstaben zu verwenden. Wenn du eine Fehlermeldung bekommst NameError: name FreqDist ist nicht definiert. Du musst deine Arbeit mit nltk. book importieren. Alle Wörter, die im letzten Beispiel produziert wurden, helfen uns, das Thema oder das Genre dieses Textes zu erfassen. Nur ein Wort, Wal. Ist etwas informativ Es kommt über 900 mal. Der Rest der Worte sagt uns nichts über den Text theyre nur englisch quotplumbing. quot Welcher Anteil des Textes mit solchen Worten aufgenommen wird, können wir eine kumulative Frequenzplot für diese Wörter mit fdist1.plot (50, cumulativeTrue) erzeugen. Um das Diagramm in 3.2 zu erzeugen. Diese 50 Wörter machen fast die Hälfte des Buches aus. Abbildung 3.2. Kumulative Frequenzplot für 50 Häufigste Wörter in Moby Dick. Diese machen fast die Hälfte der Token aus. Wenn die häufigen Worte uns nicht helfen, wie wäre es mit den Worten, die nur einmal auftreten, die sogenannten Hapaxen. Zeigen Sie sie an, indem Sie fdist1.hapaxes () eingeben. Diese Liste enthält Lexikographen. Cetologisch Schmuggelware. Expostulationen Und etwa 9.000 andere. Es scheint, dass es zu viele seltene Worte gibt, und ohne den Kontext zu sehen, können wir wohl vermuten, was die Hälfte der Hapaxen in jedem Fall bedeutet. Da weder häufige noch seltene Worte helfen, müssen wir etwas anderes ausprobieren. 3.2 Feinkörnige Wortauswahl Als nächstes schauen wir uns die langen Worte eines Textes an, die vielleicht charakteristischer und informativer werden. Dafür passen wir eine Notation von der Set-Theorie an. Wir möchten die Worte aus dem Wortschatz des Textes, die mehr als 15 Zeichen lang sind, finden. Lets nennen diese Eigenschaft P. so dass P (w) wahr ist, wenn und nur wenn w mehr als 15 Zeichen lang ist. Nun können wir die Worte von Interesse mit der mathematischen Satznotation ausdrücken, wie in (1a) gezeigt. Dies bedeutet quotthe Satz von alle w, so dass w ist ein Element von V (das Vokabular) und w hat Eigentum P quot. Daraus ergibt sich, daß die häufigste Wortlänge 3 ist und daß die Worte der Länge 3 etwa 50 000 (oder 20) der Worte ausmachen, die das Buch bilden. Obwohl wir es hier nicht weiter verfolgen werden, könnte eine weitere Analyse der Wortlänge uns dabei helfen, Unterschiede zwischen Autoren, Genres oder Sprachen zu verstehen. 3.1 fasst die in den Frequenzverteilungen definierten Funktionen zusammen. 4 Zurück zu Python: Entscheidungen treffen und kontrollieren Diese Ausdrücke haben die Form f (w) für. Oder w. f () für. . Wobei f eine Funktion ist, die mit einem Wort arbeitet, um ihre Länge zu berechnen oder sie in Großbuchstaben umzuwandeln. Jetzt müssen Sie den Unterschied zwischen den Notationen f (w) und w. f () nicht verstehen. Stattdessen lerne einfach dieses Python-Idiom, das die gleiche Operation auf jedem Element einer Liste ausführt. In den vorstehenden Beispielen geht es jedes Wort in Text1 durch. Indem man jedem wiederum die Variable w zuweist und die angegebene Operation an der Variablen ausführt. Die soeben beschriebene Notation heißt ein Quotes-Verständnis. Das ist unser erstes Beispiel für eine Python-Idiom, eine feste Notation, die wir gewöhnlich benutzen, ohne sich zu belästigen, jedes Mal zu analysieren. Das Beherrschen solcher Idiome ist ein wichtiger Teil des Fließens eines fließenden Python-Programmierers. Lass uns auf die Frage der Vokabulargröße zurückkehren und das gleiche Idiom hier anwenden: Jetzt sind wir nicht doppelzählende Wörter wie dieses und dieses. Die sich nur in der Großschreibung unterscheiden, weve um 2.000 von der Vokabeldarstellung abweichen Wir können einen Schritt weiter gehen und Zahlen und Interpunktion aus dem Vokabular zählen, indem wir alle nicht alphabetischen Items herausfiltern: Dieses Beispiel ist etwas kompliziert: Es fällt alle rein alphabetischen Items zusammen . Vielleicht wäre es einfacher gewesen, nur die Kleinbuchstaben zu zählen, aber das gibt die falsche Antwort (warum). Sorgen Sie sich nicht, wenn Sie sich nicht sicher mit Listenverständnissen fühlen, da youll viele weitere Beispiele zusammen mit Erklärungen in den folgenden Kapiteln sehen. 4.3 Verschachtelte Codeblöcke Die meisten Programmiersprachen erlauben es uns, einen Codeblock auszuführen, wenn ein bedingter Ausdruck vorliegt. Oder wenn Aussage, erfüllt ist. Wir haben bereits Beispiele von bedingten Tests in Code wie w für w in send7 gesehen, wenn len (w) lt 4. Im folgenden Programm haben wir eine Variable namens Wort mit dem String Wert cat angelegt. Die if-Anweisung prüft, ob der Test len (Wort) lt 5 wahr ist. Es ist so, dass der Körper der if-Anweisung aufgerufen wird und die print-Anweisung ausgeführt wird und eine Nachricht an den Benutzer angezeigt wird. Denken Sie daran, die Druckanweisung einzugeben, indem Sie vier Leerzeichen eingeben. Dies wird als Schleife bezeichnet, da Python den Code in kreisförmiger Weise ausführt. Es beginnt mit dem Zuweisungswort Aufruf. Effektiv die Wortvariable verwenden, um das erste Element der Liste zu nennen. Dann zeigt er den Wert des Wortes an den Benutzer an. Als nächstes geht es zurück auf die for-Anweisung und führt das Zuweisungswort mir aus. Bevor dieser neue Wert dem Benutzer angezeigt wird, und so weiter. Es geht so weiter, bis jeder Artikel der Liste verarbeitet wurde. 4.4 Schleifen mit Bedingungen Jetzt können wir die if und for-Anweisungen kombinieren. Wir werden über jedes Einzelteil der Liste schleifen und das Einzelteil nur drucken, wenn es mit dem Buchstaben l endet. Nun wählen Sie einen anderen Namen für die Variable, um zu zeigen, dass Python nicht versucht, Sinn für Variablennamen zu machen. Sie werden feststellen, dass, wenn und für Aussagen einen Doppelpunkt am Ende der Zeile haben, bevor die Einrückung beginnt. In der Tat, alle Python-Kontrollstrukturen enden mit einem Doppelpunkt. Der Doppelpunkt zeigt an, dass sich die aktuelle Aussage auf den nachfolgenden Block bezieht. Wir können auch eine Handlung angeben, die getroffen werden soll, wenn die Bedingung der if-Anweisung nicht erfüllt ist. Hier sehen wir die elif (else if) - Anweisung und die else-Anweisung. Beachten Sie, dass diese auch Colons vor dem eingerückten Code haben. 5 Automatische Natural Language Understanding Wir haben die Sprache bottom-up, mit Hilfe von Texten und der Python-Programmiersprache erforscht. Allerdings interessierten sich auch die Ausnutzung unserer Sprachkenntnisse und Berechnungen durch den Aufbau nützlicher Sprachtechnologien. Nun nehmen Sie die Gelegenheit jetzt, um von der Nitty-Kiesel des Codes zurückzukehren, um ein größeres Bild der natürlichen Sprachverarbeitung zu malen. Auf rein praktischer Ebene brauchen wir alle Hilfe, um das Universum von Informationen zu navigieren, die in Text im Web gesperrt sind. Suchmaschinen waren entscheidend für das Wachstum und die Popularität des Web, haben aber einige Mängel. Es braucht Geschick, Wissen und etwas Glück, um Antworten auf solche Fragen zu extrahieren: Welche Sehenswürdigkeiten kann ich zwischen Philadelphia und Pittsburgh auf einem begrenzten Budget besuchen Was sagen Experten über digitale SLR-Kameras Welche Vorhersagen über den Stahlmarkt wurden von glaubwürdigen gemacht Kommentatoren in der vergangenen Woche Das Erhalten eines Computers, um sie automatisch zu beantworten, beinhaltet eine Reihe von Sprachverarbeitungsaufgaben, einschließlich Informationsextraktion, Inferenz und Verdichtung, und müsste auf einer Skala und mit einer Robustheit durchgeführt werden, die noch über uns liegt Aktuelle Fähigkeiten. Auf einer philosophischeren Ebene ist eine langjährige Herausforderung in der künstlichen Intelligenz, intelligente Maschinen zu bauen, und ein großer Teil des intelligenten Verhaltens ist das Verständnis der Sprache. Seit vielen Jahren ist dieses Ziel als zu schwierig angesehen worden. Da jedoch NLP-Technologien reifer werden und robuste Methoden zur Analyse von uneingeschränktem Text zunehmend verbreitet werden, ist die Perspektive des natürlichen Sprachverständnisses als plausibles Ziel wieder aufgetaucht. In diesem Abschnitt beschreiben wir einige Sprachverständnistechnologien, um Ihnen ein Gefühl für die interessanten Herausforderungen zu geben, die auf Sie warten. 5.1 Wort Sense Disambiguation Im Wort Sinne Disambiguierung wollen wir herausfinden, welche Richtung eines Wortes in einem gegebenen Kontext beabsichtigt war. Betrachten Sie die zweideutigen Worte dienen und Gericht: dienen. Hilfe mit Essen oder Trinken halten ein Büro legen Ball in Spiel Gericht. Plattengang eines Mahlzeitkommunikationsgerätes In einem Satz, der den Ausdruck enthält: er diente dem Gericht. Sie können erkennen, dass sowohl dienen als auch Schüssel mit ihren Lebensmitteln verwendet werden. Es ist unwahrscheinlich, dass das Thema der Diskussion von Sport zu Geschirr im Raum von drei Worten verlagert wurde. Dies würde Sie zwingen, bizarre Bilder zu erfinden, wie ein Tennis-Profi, der seine oder ihre Frustrationen auf einem Porzellan-Tee-Set ausstellt, der neben dem Gericht angelegt ist. Mit anderen Worten, wir disambiguieren Worte mit Kontext, die Ausnutzung der einfachen Tatsache, dass in der Nähe Worte haben eng verwandte Bedeutungen. Als weiteres Beispiel für diese kontextuelle Wirkung, betrachten Sie das Wort von. Die mehrere Bedeutungen hat, z. B. Das Buch von Chesterton (Bevollmächtigter 8212 Chesterton war der Verfasser des Buches) die Tasse am Herd (locative 8212 der Herd ist wo die Tasse ist) und unterschreiben am Freitag (zeitlich 8212 Freitag ist die Zeit der Einreichung). Beobachten Sie in (3c), dass die Bedeutung des kursiven Wortes uns hilft, die Bedeutung von zu interpretieren. Die verlorenen Kinder wurden von den Suchern gefunden (Bevollmächtigte) Die verlorenen Kinder wurden vom Berg gefunden (locative) Die verlorenen Kinder wurden am Nachmittag (zeitlich) gefunden 5.2 Pronomen Resolution Eine tiefere Art von Sprachverständnis ist, um herauszufinden, Wann soll man die Gegenstände und Gegenstände der Verben erkennen. Du hast gelernt, dies in der Grundschule zu tun, aber es ist schwerer als du vielleicht denkst. Im Satz haben die Diebe die Gemälde gestohlen, es ist leicht zu sagen, wer die stehlende Aktion durchgeführt hat. Betrachten Sie drei mögliche Sätze in (4c). Und versuche zu bestimmen, was verkauft, gefangen und gefunden wurde (ein Fall ist zweideutig). Die Diebe stahlen die Gemälde. Sie wurden später verkauft. Die Diebe stahlen die Gemälde. Sie wurden später gefangen. Die Diebe stahlen die Gemälde. Sie wurden später gefunden. Bei der Beantwortung dieser Frage geht es darum, die Vorgänger des Pronomen zu finden. Entweder Diebe oder Gemälde. Computational Techniken zur Bewältigung dieses Problems gehören Anaphora Auflösung 8212 identifizieren, was ein Pronomen oder Nominalphrase bezieht sich auf 8212 und semantische Rollenbeschriftung 8212 identifizieren, wie eine Nominalphrase bezieht sich auf das Verb (als Agent, Patient, Instrument, und so weiter). 5.3 Erzeugen der Sprachausgabe Wenn wir diese Probleme des Sprachverständnisses automatisch lösen können, können wir auf Aufgaben umgehen, die die Erstellung von Sprachausgaben wie Fragestellung und maschinelle Übersetzung beinhalten. Im ersten Fall sollte eine Maschine in der Lage sein, auf Fragen der Texte zu antworten: Um festzustellen, ob die Hypothese durch den Text unterstützt wird, benötigt das System folgende Hintergrundwissen: (i) wenn jemand ein Autor ist Von einem Buch, dann hat er dieses Buch geschrieben (ii) wenn jemand ein Redakteur eines Buches ist, dann hat er nicht geschrieben (alle) dieses Buch (iii) wenn jemand ein Redakteur oder Autor von achtzehn Büchern ist, dann kann man nicht abschließen Das ist ein Autor von achtzehn Büchern. 5.7 Einschränkungen von NLP Trotz der forschungsorientierten Fortschritte bei Aufgaben wie RTE können natürliche Sprachsysteme, die für reale Anwendungen eingesetzt wurden, immer noch keine gesunden Menschenverstand führen oder auf weltweite Kenntnisse in einer allgemeinen und robusten Weise zurückgreifen. Wir können auf diese schwierigen künstlichen Intelligenzprobleme warten, aber in der Zwischenzeit ist es notwendig, mit einigen strengen Einschränkungen der Argumentations - und Wissensfähigkeit natürlicher Sprachsysteme zu leben. Dementsprechend war von Anfang an ein wichtiges Ziel der NLP-Forschung, Fortschritte bei der schwierigen Aufgabe zu machen, Technologien zu entwickeln, die die Sprache verstehen, indem sie oberflächliche und dennoch leistungsfähige Techniken anstelle von uneingeschränkten Kenntnissen und Argumentationsmöglichkeiten verwenden. In der Tat ist dies eines der Ziele dieses Buches, und wir hoffen, Sie mit dem Wissen und den Fähigkeiten auszustatten, um nützliche NLP-Systeme zu bauen und zur langfristigen Bestrebung des Aufbaus intelligenter Maschinen beizutragen. Texte sind in Python mit Listen dargestellt: Monty. Python Wir können Indexierung, Slicing und die Len () - Funktion auf Listen verwenden. Ein Wort quottokenquot ist ein besonderes Aussehen eines gegebenen Wortes in einem Text ein Wort quottypequot ist die einzigartige Form des Wortes als eine bestimmte Reihenfolge von Buchstaben. Wir zählen Wort-Token mit Len (Text) und Wort-Typen mit Len (Set (Text)). Wir erhalten das Vokabular eines Textes t mit sortierten (set (t)). Wir betreiben jedes Element eines Textes mit f (x) für x in Text. Um das Vokabular abzuleiten, die Untertscheidungen zu kollabieren und die Interpunktion zu ignorieren, können wir das Set (w. lower () für w in Text schreiben, wenn w. isalpha ()). Wir verarbeiten jedes Wort in einem Text mit einer for-Anweisung, z. B. für w in t: oder für Wort im Text:. Darauf folgt das Doppelpunkt und ein eingekerbter Codeblock, der jedes Mal durch die Schleife ausgeführt wird. Wir testen eine Bedingung mit einer if-Anweisung: if len (word) lt 5:. Darauf folgt das Doppelpunkt und ein eingekerbter Codeblock, der nur dann ausgeführt werden soll, wenn die Bedingung wahr ist. Eine Häufigkeitsverteilung ist eine Sammlung von Gegenständen zusammen mit ihren Frequenzzählungen (z. B. die Wörter eines Textes und ihre Häufigkeit des Aussehens). Eine Funktion ist ein Codeblock, dem ein Name zugewiesen wurde und wiederverwendet werden kann. Die Funktionen werden mit dem def-Schlüsselwort definiert, wie in def mult (x, y) x und y sind Parameter der Funktion und wirken als Platzhalter für aktuelle Datenwerte. Eine Funktion wird durch Angabe ihres Namens aufgerufen, gefolgt von Null oder mehr Argumenten in Klammern, wie folgt: Texte (). Mult (3, 4). Len (Text1). 7 Weiterführende Literatur Dieses Kapitel hat neue Konzepte in der Programmierung, der natürlichen Sprachverarbeitung und der Linguistik eingeführt, die alle zusammen gemischt sind. Viele von ihnen werden in den folgenden Kapiteln konsolidiert. Sie können aber auch die mit diesem Kapitel gelieferten Online-Materialien (bei nltk. org), einschließlich Links zu zusätzlichen Hintergrundmaterialien und Links zu Online-NLP-Systemen, konsultieren. Sie können auch gerne auf einige Linguistik und NLP-bezogene Konzepte in Wikipedia (z. B. Kollokationen, die Turing-Test, die Typ-Token-Unterscheidung) zu lesen. Sie sollten sich mit der Python-Dokumentation bei docs. python. org vertraut machen. Einschließlich der vielen Tutorials und umfassenden Referenzmaterialien, die dort verlinkt sind. Ein Anfänger-Führer zu Python ist bei wiki. python. orgmoinBeginnersGuide verfügbar. Verschiedene Fragen über Python könnten in der FAQ bei python. orgdocfaqgeneral beantwortet werden. Wie Sie in NLTK vertiefen, möchten Sie vielleicht die Mailing-Liste abonnieren, wo neue Releases des Toolkits angekündigt werden. Es gibt auch eine NLTK-User Mailing-Liste, wo Benutzer einander helfen, wie sie lernen, wie man Python und NLTK für Sprachanalyse verwenden. Details zu diesen Listen finden Sie unter nltk. org. Für weitere Informationen zu den Themen, die in 5. und auf NLP allgemeiner gesprochen werden, können Sie gerne eines der folgenden ausgezeichneten Bücher konsultieren: Indurkhya, Nitin und Fred Damerau (Hrsg., 2010) Handbuch der natürlichen Sprachverarbeitung (zweite Auflage) Chapman amp HallCRC 2010. (Indurkhya amp Damerau, 2010) (Dale, Moisl, amp Somers, 2000) Jurafsky, Daniel und James Martin (2008) Sprach - und Sprachverarbeitung (Zweite Auflage). Prentice Hall. (Jurafsky amp Martin, 2008) Mitkov, Ruslan (Hrsg. 2003) Das Oxford-Handbuch der Computational Linguistics. Oxford University Press. (Zweite Auflage 2010 erwartet). (Mitkov, 2002) Die Vereinigung für Computational Linguistics ist die internationale Organisation, die das Feld der NLP darstellt. Die ACL-Website (aclweb. org) beherbergt viele nützliche Ressourcen, darunter: Informationen über internationale und regionale Konferenzen und Workshops das ACL Wiki mit Links zu Hunderten von nützlichen Ressourcen und der ACL Anthology. Die die meisten der NLP-Forschungsliteratur aus den vergangenen 50 Jahren enthält, vollständig indiziert und frei herunterladbar. Einige ausgezeichnete einführende Linguistik Lehrbücher sind: Finegan2007. (OGrady et al., 2004). (OSU, 2007). Vielleicht möchten Sie auch LanguageLog konsultieren. Ein populärer Linguistik-Blog mit gelegentlichen Beiträgen, die die in diesem Buch beschriebenen Techniken verwenden. 8 Übungen 9788 Versuchen Sie es mit dem Python-Interpreter als Taschenrechner und geben Sie Ausdrücke wie 12 (4 1) ein. 9788 Bei einem Alphabet von 26 Briefen gibt es 26 an die Macht 10 oder 26 10. Zehn-Buchstaben-Strings können wir bilden. Das klappt aus 141167095653376 Wie viele Hundert-Buchstaben-Strings sind möglich 9788 Die Python-Multiplikationsoperation kann auf Listen angewendet werden. Was passiert, wenn Sie Monty eingeben. Python 20 Oder 3 sent1 9788 Review 1 auf Computing mit Sprache. Wie viele Wörter gibt es in Text2. Wie viele deutliche Worte gibt es 9788 Vergleichen Sie die lexikalischen Vielfalt Scores für Humor und Romantik Fiktion in 1.1. Welches Genre ist mehr lexikalisch vielfältig 9788 Produzieren Sie eine Dispersionsdarstellung der vier Hauptprotagonisten in Sinn und Sinnlichkeit. Elinor, Marianne, Edward und Willoughby. Was können Sie über die verschiedenen Rollen von den Männern und Frauen in diesem Roman zu sehen, können Sie die Paare identifizieren 9788 Finden Sie die Kollokationen in Text5. 9788 Betrachten Sie den folgenden Python-Ausdruck: len (set (text4)). Geben Sie den Zweck dieses Ausdrucks an. Beschreibe die beiden Schritte bei der Durchführung dieser Berechnung. 9788 Review 2 auf Listen und Strings. Definiere einen String und ordne ihn einer Variablen zu, z. B. Mystring My String (aber etwas interessanter in der Saite). Drucken Sie den Inhalt dieser Variablen auf zwei Arten, indem Sie einfach den Variablennamen eingeben und die Eingabetaste drücken und dann die Druckanweisung verwenden. Versuchen Sie, die String zu sich selbst mit Mystring Mystring. Oder multipliziert mit einer Zahl, z. B. Mystring 3 Beachten Sie, dass die Saiten ohne Leerzeichen miteinander verbunden sind. Wie könnte man diese 9788 definieren Definiere eine Variable mysent, um eine Liste von Wörtern zu sein, mit der Syntax mysent quotMyquot. Quoteentquot (aber mit deinen eigenen Worten oder einem Lieblings-Sprichwort). Verwenden Sie. join (mysent), um diese in einen String zu konvertieren. Verwenden Sie split (), um den String wieder in das Listenformular aufzuteilen, mit dem Sie beginnen mussten. 9788 Definieren Sie mehrere Variablen mit Listen von Wörtern, z. B. Phrase1 Phrase2 und so weiter. Verbinden Sie sie zusammen in verschiedenen Kombinationen (mit dem Plus-Operator), um ganze Sätze zu bilden. Was ist die Beziehung zwischen len (phrase1 phrase2) und len (phrase1) len (phrase2) 9788 Betrachten Sie die folgenden zwei Ausdrücke, die den gleichen Wert haben. Which one will typically be more relevant in NLP Why 9788 We have seen how to represent a sentence as a list of words, where each word is a sequence of characters. What does sent122 do Why Experiment with other index values. 9788 The first sentence of text3 is provided to you in the variable sent3 . The index of the in sent3 is 1, because sent31 gives us the . What are the indexes of the two other occurrences of this word in sent3 9788 Review the discussion of conditionals in 4. Find all words in the Chat Corpus ( text5 ) starting with the letter b. Show them in alphabetical order. 9788 Type the expression list(range(10)) at the interpreter prompt. Now try list(range(10, 20)) . list(range(10, 20, 2)) . and list(range(20, 10, -2)) . We will see a variety of uses for this built-in function in later chapters. 9681 Use text9.index() to find the index of the word sunset. Youll need to insert this word as an argument between the parentheses. By a process of trial and error, find the slice for the complete sentence that contains this word. 9681 Using list addition, and the set and sorted operations, compute the vocabulary of the sentences sent1 . sent8 . 9681 What is the difference between the following two lines Which one will give a larger value Will this be the case for other texts Docutils System Messages1 Language Processing and Python It is easy to get our hands on millions of words of text. What can we do with it, assuming we can write some simple programs In this chapter well address the following questions: What can we achieve by combining simple programming techniques with large quantities of text How can we automatically extract key words and phrases that sum up the style and content of a text What tools and techniques does the Python programming language provide for such work What are some of the interesting challenges of natural language processing This chapter is divided into sections that skip between two quite different styles. In the quotcomputing with languagequot sections we will take on some linguistically motivated programming tasks without necessarily explaining how they work. In the quotcloser look at Pythonquot sections we will systematically review key programming concepts. Well flag the two styles in the section titles, but later chapters will mix both styles without being so up-front about it. We hope this style of introduction gives you an authentic taste of what will come later, while covering a range of elementary concepts in linguistics and computer science. If you have basic familiarity with both areas, you can skip to 1.5 we will repeat any important points in later chapters, and if you miss anything you can easily consult the online reference material at nltk. org . If the material is completely new to you, this chapter will raise more questions than it answers, questions that are addressed in the rest of this book. 1.1 Computing with Language: Texts and Words Were all very familiar with text, since we read and write it every day. Here we will treat text as raw data for the programs we write, programs that manipulate and analyze it in a variety of interesting ways. But before we can do this, we have to get started with the Python interpreter. Getting Started with Python One of the friendly things about Python is that it allows you to type directly into the interactive interpreter 8212 the program that will be running your Python programs. You can access the Python interpreter using a simple graphical interface called the Interactive DeveLopment Environment (IDLE). On a Mac you can find this under Applications 8594 MacPython . and on Windows under All Programs 8594 Python . Under Unix you can run Python from the shell by typing idle (if this is not installed, try typing python ). The interpreter will print a blurb about your Python version simply check that you are running Python 2.4 or 2.5 (here it is 2.5.1): Your Turn: Try searching for other words to save re-typing, you might be able to use up-arrow, Ctrl-up-arrow or Alt-p to access the previous command and modify the word being searched. You can also try searches on some of the other texts we have included. For example, search Sense and Sensibility for the word affection. using text2.concordance( quotaffectionquot ) . Search the book of Genesis to find out how long some people lived, using text3.concordance( quotlivedquot ) . You could look at text4 . the Inaugural Address Corpus . to see examples of English going back to 1789, and search for words like nation. terror. god to see how these words have been used differently over time. Weve also included text5 . the NPS Chat Corpus . search this for unconventional words like im. ur. Lol (Note that this corpus is uncensored) Once youve spent a little while examining these texts, we hope you have a new sense of the richness and diversity of language. In the next chapter you will learn how to access a broader range of text, including text in languages other than English. A concordance permits us to see words in context. For example, we saw that monstrous occurred in contexts such as the pictures and the size. What other words appear in a similar range of contexts We can find out by appending the term similar to the name of the text in question, then inserting the relevant word in parentheses: Observe that we get different results for different texts. Austen uses this word quite differently from Melville for her, monstrous has positive connotations, and sometimes functions as an intensifier like the word very . The term commoncontexts allows us to examine just the contexts that are shared by two or more words, such as monstrous and very. We have to enclose these words by square brackets as well as parentheses, and separate them with a comma: Your Turn: Pick another pair of words and compare their usage in two different texts, using the similar() and commoncontexts() functions. It is one thing to automatically detect that a particular word occurs in a text, and to display some words that appear in the same context. However, we can also determine the location of a word in the text: how many words from the beginning it appears. This positional information can be displayed using a dispersion plot. Each stripe represents an instance of a word, and each row represents the entire text. In 1.2 we see some striking patterns of word usage over the last 220 years (in an artificial text constructed by joining the texts of the Inaugural Address Corpus end-to-end). You can produce this plot as shown below. You might like to try more words (e. g. liberty. constitution ), and different texts. Can you predict the dispersion of a word before you view it As before, take care to get the quotes, commas, brackets and parentheses exactly right. Figure 1.2. Lexical Dispersion Plot for Words in U. S. Presidential Inaugural Addresses: This can be used to investigate changes in language use over time. Important: You need to have Pythons NumPy and Matplotlib packages installed in order to produce the graphical plots used in this book. Please see nltk. org for installation instructions. Now, just for fun, lets try generating some random text in the various styles we have just seen. To do this, we type the name of the text followed by the term generate . (We need to include the parentheses, but theres nothing that goes between them.) Note that the first time you run this command, it is slow because it gathers statistics about word sequences. Each time you run it, you will get different output text. Now try generating random text in the style of an inaugural address or an Internet chat room. Although the text is random, it re-uses common words and phrases from the source text and gives us a sense of its style and content. (What is lacking in this randomly generated text) When generate produces its output, punctuation is split off from the preceding word. While this is not correct formatting for English text, we do it to make clear that words and punctuation are independent of one another. You will learn more about this in 3 . Counting Vocabulary The most obvious fact about texts that emerges from the preceding examples is that they differ in the vocabulary they use. In this section we will see how to use the computer to count the words in a text in a variety of useful ways. As before, you will jump right in and experiment with the Python interpreter, even though you may not have studied Python systematically yet. Test your understanding by modifying the examples, and trying the exercises at the end of the chapter. Lets begin by finding out the length of a text from start to finish, in terms of the words and punctuation symbols that appear. We use the term len to get the length of something, which well apply here to the book of Genesis: So Genesis has 44,764 words and punctuation symbols, or quottokens. quot A token is the technical name for a sequence of characters 8212 such as hairy . his . or :) 8212 that we want to treat as a group. When we count the number of tokens in a text, say, the phrase to be or not to be. we are counting occurrences of these sequences. Thus, in our example phrase there are two occurrences of to. two of be. and one each of or and not. But there are only four distinct vocabulary items in this phrase. How many distinct words does the book of Genesis contain To work this out in Python, we have to pose the question slightly differently. The vocabulary of a text is just the set of tokens that it uses, since in a set, all duplicates are collapsed together. In Python we can obtain the vocabulary items of text3 with the command: set(text3) . When you do this, many screens of words will fly past. Now try the following: By wrapping sorted() around the Python expression set(text3) . we obtain a sorted list of vocabulary items, beginning with various punctuation symbols and continuing with words starting with A. All capitalized words precede lowercase words. We discover the size of the vocabulary indirectly, by asking for the number of items in the set, and again we can use len to obtain this number . Although it has 44,764 tokens, this book has only 2,789 distinct words, or quotword types. quot A word type is the form or spelling of the word independently of its specific occurrences in a text 8212 that is, the word considered as a unique item of vocabulary. Our count of 2,789 items will include punctuation symbols, so we will generally call these unique items types instead of word types. Now, lets calculate a measure of the lexical richness of the text. The next example shows us that each word is used 16 times on average (we need to make sure Python uses floating point division): Notice that our indexes start from zero: sent element zero, written sent0 . is the first word, word1 . whereas sent element 9 is word10 . The reason is simple: the moment Python accesses the content of a list from the computers memory, it is already at the first element we have to tell it how many elements forward to go. Thus, zero steps forward leaves it at the first element. This practice of counting from zero is initially confusing, but typical of modern programming languages. Youll quickly get the hang of it if youve mastered the system of counting centuries where 19XY is a year in the 20th century, or if you live in a country where the floors of a building are numbered from 1, and so walking up n-1 flights of stairs takes you to level n . Now, if we accidentally use an index that is too large, we get an error: This time it is not a syntax error, because the program fragment is syntactically correct. Instead, it is a runtime error. and it produces a Traceback message that shows the context of the error, followed by the name of the error, IndexError . and a brief explanation. Lets take a closer look at slicing, using our artificial sentence again. Here we verify that the slice 5:8 includes sent elements at indexes 5, 6, and 7: We can modify an element of a list by assigning to one of its index values. In the next example, we put sent0 on the left of the equals sign . We can also replace an entire slice with new material . A consequence of this last change is that the list only has four elements, and accessing a later value generates an error . Your Turn: Take a few minutes to define a sentence of your own and modify individual words and groups of words (slices) using the same methods used earlier. Check your understanding by trying the exercises on lists at the end of this chapter. From the start of 1.1. you have had access to texts called text1 . text2 . und so weiter. It saved a lot of typing to be able to refer to a 250,000-word book with a short name like this In general, we can make up names for anything we care to calculate. We did this ourselves in the previous sections, e. g. defining a variable sent1 . as follows: Such lines have the form: variable expression . Python will evaluate the expression, and save its result to the variable. This process is called assignment. It does not generate any output you have to type the variable on a line of its own to inspect its contents. The equals sign is slightly misleading, since information is moving from the right side to the left. It might help to think of it as a left-arrow. The name of the variable can be anything you like, e. g. mysent . sentence . xyzzy . It must start with a letter, and can include numbers and underscores. Here are some examples of variables and assignments: Remember that capitalized words appear before lowercase words in sorted lists. Notice in the previous example that we split the definition of mysent over two lines. Python expressions can be split across multiple lines, so long as this happens within any kind of brackets. Python uses the quot . quot prompt to indicate that more input is expected. It doesnt matter how much indentation is used in these continuation lines, but some indentation usually makes them easier to read. It is good to choose meaningful variable names to remind you 8212 and to help anyone else who reads your Python code 8212 what your code is meant to do. Python does not try to make sense of the names it blindly follows your instructions, and does not object if you do something confusing, such as one two or two 3 . The only restriction is that a variable name cannot be any of Pythons reserved words, such as def . if . not . and import . If you use a reserved word, Python will produce a syntax error: Take care with your choice of names (or identifiers ) for Python variables. First, you should start the name with a letter, optionally followed by digits ( 0 to 9 ) or letters. Thus, abc23 is fine, but 23abc will cause a syntax error. Names are case-sensitive, which means that myVar and myvar are distinct variables. Variable names cannot contain whitespace, but you can separate words using an underscore, e. g. myvar . Be careful not to insert a hyphen instead of an underscore: my-var is wrong, since Python interprets the quot - quot as a minus sign. Some of the methods we used to access the elements of a list also work with individual words, or strings. For example, we can assign a string to a variable . index a string . and slice a string : We can also perform multiplication and addition with strings: We can join the words of a list to make a single string, or split a string into a list, as follows: We will come back to the topic of strings in 3. For the time being, we have two important building blocks 8212 lists and strings 8212 and are ready to get back to some language analysis. 1.3 Computing with Language: Simple Statistics Lets return to our exploration of the ways we can bring our computational resources to bear on large quantities of text. We began this discussion in 1.1. and saw how to search for words in context, how to compile the vocabulary of a text, how to generate random text in the same style, and so on. In this section we pick up the question of what makes a text distinct, and use automatic methods to find characteristic words and expressions of a text. As in 1.1. you can try new features of the Python language by copying them into the interpreter, and youll learn about these features systematically in the following section. Before continuing further, you might like to check your understanding of the last section by predicting the output of the following code. You can use the interpreter to check whether you got it right. If youre not sure how to do this task, it would be a good idea to review the previous section before continuing further. Frequency Distributions How can we automatically identify the words of a text that are most informative about the topic and genre of the text Imagine how you might go about finding the 50 most frequent words of a book. One method would be to keep a tally for each vocabulary item, like that shown in 1.3. The tally would need thousands of rows, and it would be an exceedingly laborious process 8212 so laborious that we would rather assign the task to a machine. Figure 1.3. Counting Words Appearing in a Text (a frequency distribution) The table in 1.3 is known as a frequency distribution. and it tells us the frequency of each vocabulary item in the text. (In general, it could count any kind of observable event.) It is a quotdistributionquot because it tells us how the total number of word tokens in the text are distributed across the vocabulary items. Since we often need frequency distributions in language processing, NLTK provides built-in support for them. Lets use a FreqDist to find the 50 most frequent words of Moby Dick . Try to work out what is going on here, then read the explanation that follows. When we first invoke FreqDist . we pass the name of the text as an argument . We can inspect the total number of words (quotoutcomesquot) that have been counted up 8212 260,819 in the case of Moby Dick . The expression keys () gives us a list of all the distinct types in the text . and we can look at the first 50 of these by slicing the list . Your Turn: Try the preceding frequency distribution example for yourself, for text2 . Be careful to use the correct parentheses and uppercase letters. If you get an error message NameError: name FreqDist is not defined . you need to start your work with from nltk. book import Do any words produced in the last example help us grasp the topic or genre of this text Only one word, whale. is slightly informative It occurs over 900 times. The rest of the words tell us nothing about the text theyre just English quotplumbing. quot What proportion of the text is taken up with such words We can generate a cumulative frequency plot for these words, using fdist1.plot(50, cumulativeTrue) . to produce the graph in 1.4. These 50 words account for nearly half the book Figure 1.4. Cumulative Frequency Plot for 50 Most Frequently Words in Moby Dick . these account for nearly half of the tokens. If the frequent words dont help us, how about the words that occur once only, the so-called hapaxes. View them by typing fdist1.hapaxes() . This list contains lexicographer. cetological. contraband. expostulations. and about 9,000 others. It seems that there are too many rare words, and without seeing the context we probably cant guess what half of the hapaxes mean in any case Since neither frequent nor infrequent words help, we need to try something else. Fine-grained Selection of Words Next, lets look at the long words of a text perhaps these will be more characteristic and informative. For this we adapt some notation from set theory. We would like to find the words from the vocabulary of the text that are more than 15 characters long. Lets call this property P. so that P(w) is true if and only if w is more than 15 characters long. Now we can express the words of interest using mathematical set notation as shown in (1a). This means quotthe set of all w such that w is an element of V (the vocabulary) and w has property P quot. From this we see that the most frequent word length is 3, and that words of length 3 account for roughly 50,000 (or 20) of the words making up the book. Although we will not pursue it here, further analysis of word length might help us understand differences between authors, genres, or languages. 1.2 summarizes the functions defined in frequency distributions. These expressions have the form f(w) for. or w. f() for. . where f is a function that operates on a word to compute its length, or to convert it to uppercase. For now, you dont need to understand the difference between the notations f(w) and w. f() . Instead, simply learn this Python idiom which performs the same operation on every element of a list. In the preceding examples, it goes through each word in text1 . assigning each one in turn to the variable w and performing the specified operation on the variable. The notation just described is called a quotlist comprehension. quot This is our first example of a Python idiom, a fixed notation that we use habitually without bothering to analyze each time. Mastering such idioms is an important part of becoming a fluent Python programmer. Lets return to the question of vocabulary size, and apply the same idiom here: Now that we are not double-counting words like This and this. which differ only in capitalization, weve wiped 2,000 off the vocabulary count We can go a step further and eliminate numbers and punctuation from the vocabulary count by filtering out any non-alphabetic items: This example is slightly complicated: it lowercases all the purely alphabetic items. Perhaps it would have been simpler just to count the lowercase-only items, but this gives the wrong answer (why). Dont worry if you dont feel confident with list comprehensions yet, since youll see many more examples along with explanations in the following chapters. Nested Code Blocks Most programming languages permit us to execute a block of code when a conditional expression. or if statement, is satisfied. We already saw examples of conditional tests in code like w for w in sent7 if len(w) lt 4 . In the following program, we have created a variable called word containing the string value cat . The if statement checks whether the test len(word) lt 5 is true. It is, so the body of the if statement is invoked and the print statement is executed, displaying a message to the user. Remember to indent the print statement by typing four spaces. This is called a loop because Python executes the code in circular fashion. It starts by performing the assignment word Call . effectively using the word variable to name the first item of the list. Then, it displays the value of word to the user. Next, it goes back to the for statement, and performs the assignment word me . before displaying this new value to the user, and so on. It continues in this fashion until every item of the list has been processed. Looping with Conditions Now we can combine the if and for statements. We will loop over every item of the list, and print the item only if it ends with the letter l . Well pick another name for the variable to demonstrate that Python doesnt try to make sense of variable names. You will notice that if and for statements have a colon at the end of the line, before the indentation begins. In fact, all Python control structures end with a colon. The colon indicates that the current statement relates to the indented block that follows. We can also specify an action to be taken if the condition of the if statement is not met. Here we see the elif (else if) statement, and the else statement. Notice that these also have colons before the indented code. 1.5 Automatic Natural Language Understanding We have been exploring language bottom-up, with the help of texts and the Python programming language. However, were also interested in exploiting our knowledge of language and computation by building useful language technologies. Well take the opportunity now to step back from the nitty-gritty of code in order to paint a bigger picture of natural language processing. At a purely practical level, we all need help to navigate the universe of information locked up in text on the Web. Search engines have been crucial to the growth and popularity of the Web, but have some shortcomings. It takes skill, knowledge, and some luck, to extract answers to such questions as: What tourist sites can I visit between Philadelphia and Pittsburgh on a limited budget What do experts say about digital SLR cameras What predictions about the steel market were made by credible commentators in the past week Getting a computer to answer them automatically involves a range of language processing tasks, including information extraction, inference, and summarization, and would need to be carried out on a scale and with a level of robustness that is still beyond our current capabilities. On a more philosophical level, a long-standing challenge within artificial intelligence has been to build intelligent machines, and a major part of intelligent behaviour is understanding language. For many years this goal has been seen as too difficult. However, as NLP technologies become more mature, and robust methods for analyzing unrestricted text become more widespread, the prospect of natural language understanding has re-emerged as a plausible goal. In this section we describe some language understanding technologies, to give you a sense of the interesting challenges that are waiting for you. Word Sense Disambiguation In word sense disambiguation we want to work out which sense of a word was intended in a given context. Consider the ambiguous words serve and dish : serve. help with food or drink hold an office put ball into play dish. plate course of a meal communications device In a sentence containing the phrase: he served the dish. you can detect that both serve and dish are being used with their food meanings. Its unlikely that the topic of discussion shifted from sports to crockery in the space of three words. This would force you to invent bizarre images, like a tennis pro taking out their frustrations on a china tea-set laid out beside the court. In other words, we automatically disambiguate words using context, exploiting the simple fact that nearby words have closely related meanings. As another example of this contextual effect, consider the word by. which has several meanings, e. g. the book by Chesterton (agentive 8212 Chesterton was the author of the book) the cup by the stove (locative 8212 the stove is where the cup is) and submit by Friday (temporal 8212 Friday is the time of the submitting). Observe in (3c) that the meaning of the italicized word helps us interpret the meaning of by . The lost children were found by the searchers (agentive) The lost children were found by the mountain (locative) The lost children were found by the afternoon (temporal) Pronoun Resolution A deeper kind of language understanding is to work out quotwho did what to whomquot 8212 i. e. to detect the subjects and objects of verbs. You learnt to do this in elementary school, but its harder than you might think. In the sentence the thieves stole the paintings it is easy to tell who performed the stealing action. Consider three possible following sentences in (4c). and try to determine what was sold, caught, and found (one case is ambiguous). The thieves stole the paintings. They were subsequently sold . The thieves stole the paintings. They were subsequently caught . The thieves stole the paintings. They were subsequently found . Answering this question involves finding the antecedent of the pronoun they. either thieves or paintings. Computational techniques for tackling this problem include anaphora resolution 8212 identifying what a pronoun or noun phrase refers to 8212 and semantic role labeling 8212 identifying how a noun phrase relates to the verb (as agent, patient, instrument, and so on). Generating Language Output If we can automatically solve such problems of language understanding, we will be able to move on to tasks that involve generating language output, such as question answering and machine translation. In the first case, a machine should be able to answer a users questions relating to collection of texts: In order to determine whether the hypothesis is supported by the text, the system needs the following background knowledge: (i) if someone is an author of a book, then heshe has written that book (ii) if someone is an editor of a book, then heshe has not written (all of) that book (iii) if someone is editor or author of eighteen books, then one cannot conclude that heshe is author of eighteen books. Limitations of NLP Despite the research-led advances in tasks like RTE, natural language systems that have been deployed for real-world applications still cannot perform common-sense reasoning or draw on world knowledge in a general and robust manner. We can wait for these difficult artificial intelligence problems to be solved, but in the meantime it is necessary to live with some severe limitations on the reasoning and knowledge capabilities of natural language systems. Accordingly, right from the beginning, an important goal of NLP research has been to make progress on the difficult task of building technologies that quotunderstand language, quot using superficial yet powerful techniques instead of unrestricted knowledge and reasoning capabilities. Indeed, this is one of the goals of this book, and we hope to equip you with the knowledge and skills to build useful NLP systems, and to contribute to the long-term aspiration of building intelligent machines. 1.6 Summary Texts are represented in Python using lists: Monty. Python . We can use indexing, slicing, and the len() function on lists. A word quottokenquot is a particular appearance of a given word in a text a word quottypequot is the unique form of the word as a particular sequence of letters. We count word tokens using len(text) and word types using len(set(text)) . We obtain the vocabulary of a text t using sorted(set(t)) . We operate on each item of a text using f(x) for x in text . To derive the vocabulary, collapsing case distinctions and ignoring punctuation, we can write set(w. lower() for w in text if w. isalpha()) . We process each word in a text using a for statement, such as for w in t: or for word in text: . This must be followed by the colon character and an indented block of code, to be executed each time through the loop. We test a condition using an if statement: if len(word) lt 5: . This must be followed by the colon character and an indented block of code, to be executed only if the condition is true. A frequency distribution is a collection of items along with their frequency counts (e. g. the words of a text and their frequency of appearance). A function is a block of code that has been assigned a name and can be reused. Functions are defined using the def keyword, as in def mult (x, y) x and y are parameters of the function, and act as placeholders for actual data values. A function is called by specifying its name followed by one or more arguments inside parentheses, like this: mult(3, 4) . z. B. len(text1) . 1.7 Further Reading This chapter has introduced new concepts in programming, natural language processing, and linguistics, all mixed in together. Many of them are consolidated in the following chapters. However, you may also want to consult the online materials provided with this chapter (at nltk. org ), including links to additional background materials, and links to online NLP systems. You may also like to read up on some linguistics and NLP-related concepts in Wikipedia (e. g. collocations, the Turing Test, the type-token distinction). You should acquaint yourself with the Python documentation available at docs. python. org . including the many tutorials and comprehensive reference materials linked there. A Beginners Guide to Python is available at wiki. python. orgmoinBeginnersGuide . Miscellaneous questions about Python might be answered in the FAQ at python. orgdocfaqgeneral . As you delve into NLTK, you might want to subscribe to the mailing list where new releases of the toolkit are announced. There is also an NLTK-Users mailing list, where users help each other as they learn how to use Python and NLTK for language analysis work. Details of these lists are available at nltk. org . For more information on the topics covered in 1.5. and on NLP more generally, you might like to consult one of the following excellent books: Indurkhya, Nitin and Fred Damerau (eds, 2010) Handbook of Natural Language Processing (Second Edition) Chapman amp HallCRC. 2010. (Indurkhya amp Damerau, 2010) (Dale, Moisl, amp Somers, 2000) Jurafsky, Daniel and James Martin (2008) Speech and Language Processing (Second Edition). Prentice Hall. (Jurafsky amp Martin, 2008) Mitkov, Ruslan (ed, 2003) The Oxford Handbook of Computational Linguistics . Oxford University Press. (second edition expected in 2010). (Mitkov, 2002) The Association for Computational Linguistics is the international organization that represents the field of NLP. The ACL website ( aclweb. org ) hosts many useful resources, including: information about international and regional conferences and workshops the ACL Wiki with links to hundreds of useful resources and the ACL Anthology. which contains most of the NLP research literature from the past 50 years, fully indexed and freely downloadable. Some excellent introductory Linguistics textbooks are: (Finegan, 2007). (OGrady et al, 2004). (OSU, 2007). You might like to consult LanguageLog. a popular linguistics blog with occasional posts that use the techniques described in this book. 1.8 Exercises 9788 Try using the Python interpreter as a calculator, and typing expressions like 12 (4 1) . 9788 Given an alphabet of 26 letters, there are 26 to the power 10, or 26 10 . ten-letter strings we can form. That works out to 141167095653376L (the L at the end just indicates that this is Pythons long-number format). How many hundred-letter strings are possible 9788 The Python multiplication operation can be applied to lists. What happens when you type Monty. Python 20 . or 3 sent1 9788 Review 1.1 on computing with language. How many words are there in text2 . How many distinct words are there 9788 Compare the lexical diversity scores for humor and romance fiction in 1.1. Which genre is more lexically diverse 9788 Produce a dispersion plot of the four main protagonists in Sense and Sensibility . Elinor, Marianne, Edward, and Willoughby. What can you observe about the different roles played by the males and females in this novel Can you identify the couples 9788 Find the collocations in text5 . 9788 Consider the following Python expression: len(set(text4)) . State the purpose of this expression. Describe the two steps involved in performing this computation. 9788 Review 1.2 on lists and strings. Define a string and assign it to a variable, e. g. mystring My String (but put something more interesting in the string). Print the contents of this variable in two ways, first by simply typing the variable name and pressing enter, then by using the print statement. Try adding the string to itself using mystring mystring . or multiplying it by a number, e. g. mystring 3 . Notice that the strings are joined together without any spaces. How could you fix this 9788 Define a variable mysent to be a list of words, using the syntax mysent quotMyquot. quotsentquot (but with your own words, or a favorite saying). Use. join(mysent) to convert this into a string. Use split() to split the string back into the list form you had to start with. 9788 Define several variables containing lists of words, e. g. phrase1 . phrase2 . und so weiter. Join them together in various combinations (using the plus operator) to form whole sentences. What is the relationship between len(phrase1 phrase2) and len(phrase1) len(phrase2) 9788 Consider the following two expressions, which have the same value. Which one will typically be more relevant in NLP Why 9788 We have seen how to represent a sentence as a list of words, where each word is a sequence of characters. What does sent122 do Why Experiment with other index values. 9788 The first sentence of text3 is provided to you in the variable sent3 . The index of the in sent3 is 1, because sent31 gives us the . What are the indexes of the two other occurrences of this word in sent3 9788 Review the discussion of conditionals in 1.4. Find all words in the Chat Corpus ( text5 ) starting with the letter b. Show them in alphabetical order. 9788 Type the expression range(10) at the interpreter prompt. Now try range(10, 20) . range(10, 20, 2) . and range(20, 10, -2) . We will see a variety of uses for this built-in function in later chapters. 9681 Use text9.index() to find the index of the word sunset. Youll need to insert this word as an argument between the parentheses. By a process of trial and error, find the slice for the complete sentence that contains this word. 9681 Using list addition, and the set and sorted operations, compute the vocabulary of the sentences sent1 . sent8 . 9681 What is the difference between the following two lines Which one will give a larger value Will this be the case for other texts
NN, Inc. ist ein diversifiziertes Industrieunternehmen und ein weltweit führender Hersteller von hochpräzisen Lagerkomponenten, industriellen Kunststoffprodukten und Präzisionsmetallkomponenten für eine Vielzahl von Märkten. NN, Inc. ist ein diversifiziertes Industrieunternehmen und ein weltweit führender Hersteller von hochpräzisen Lagerkomponenten, Auf globaler Basis. Wir haben 42 Produktionsstätten in Nordamerika, Westeuropa, Osteuropa, Südamerika und China. Wie in diesem Geschäftsbericht auf Formular 10-K verwendet, sind die Bedingungen NN, die Gesellschaft, wir, unsere oder wir NN, Inc. und ihre Tochtergesellschaften. Unser Geschäft ist in drei berichtspflichtige Segmente, die Precision Bearing Components Group (früher bekannt als unsere Metal Bearing Components Group), die Precision Engineered Products Group (früher bekannt als unsere Plastics and Rubber Components Group) und die Autocam Precision Components Group zusammengefasst. Wir haben im Jahr 2015 zwei Akquisitionen und ein...
Comments
Post a Comment