Natural Language Processing Lab guide
Guide to NLP Lab and its codes
We were given the problems for the first part of the natural language processing lab work as,
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
Lab problems to solve tomorrow and Next classes
1. Write a Python Program to perform following tasks on text
"JKKNIU stands as a dedicated center for higher education, honoring the profound cultural and literary legacy of Bangladesh's national poet."
a) Tokenization
b) Stop word Removal
2. Write Python Program for
a) Word Analysis
b) Word Generation.
3. Create a Sample list for at least 5 words with ambiguous sense and Write a Python program to implement WSD.
4. Install NLTK tool kit and perform stemming.
5. Create Sample list of at least 10 words POS tagging and find the POS for any given word.
6. Write a Python program to
a) Perform Morphological Analysis using NLTK library
b) Generate n-grams using NLTK N-Grams library
c) Implement N-Grams Smoothing.
7. Using NLTK package to convert audio file to text and text file to audio files.
Text: "Jatiya Kabi Kazi Nazrul Islam University (JKKNIU) is a premier public university in Trishal, Mymensingh, established in 2006."
before we work with the code here are the libraries that are needed.
1
pip install nltk gTTS pydub SpeechRecognition
So we have to use python for the tasks and we have to use NLTK library for the tasks, and write the codes as well.
1. Tokenization + Stop word removal
The first task we are given is as follows,
- Write a Python Program to perform following tasks on text “JKKNIU stands as a dedicated center for higher education, honoring the profound cultural and literary legacy of Bangladesh’s national poet.”
- Tokenization
- Stop word Removal
And the code for this is,
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
from nltk.tokenize import word_tokenize, sent_tokenize
from nltk.corpus import stopwords
nltk.download('punkt')
nltk.download('punkt_tab')
nltk.download('stopwords')
text = "JKKNIU stands as a dedicated center for higher education, honoring the profound cultural and literary legacy of Bangladesh's national poet."
tokens = word_tokenize(text)
print("Tokens:", tokens)
print("Sentences:", sent_tokenize(text))
stop = set(stopwords.words('english'))
filtered = [w for w in tokens if w.lower() not in stop and w.isalnum()]
print("After stop word removal:", filtered)
here we have to use the punkt and punkt_tab for tokenization and stopwords for stop word removal.
Stemming
In the second problem we are asked to
- Write Python Program for,
- Word Analysis
- Word Generation.
so the code is
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
import nltk
from nltk.tokenize import word_tokenize, sent_tokenize
from nltk.corpus import stopwords, wordnet
from nltk import FreqDist
from nltk.stem import PorterStemmer, WordNetLemmatizer
nltk.download('punkt')
nltk.download('punkt_tab')
nltk.download('wordnet')
text = "JKKNIU stands as a dedicated center for higher education, honoring the profound cultural and literary legacy of Bangladesh's national poet."
tokens = word_tokenize(text)
# a) Word analysis: frequency, length, root and affix
words = [w.lower() for w in tokens if w.isalpha()]
print("Total:", len(words), "| Unique:", len(set(words)))
print("Most common:", FreqDist(words).most_common(3))
ps, wl = PorterStemmer(), WordNetLemmatizer()
for w in ['honoring', 'dedicated', 'stands']:
stem = ps.stem(w)
print(w, "| length:", len(w), "| stem:", stem,
"| lemma:", wl.lemmatize(w, 'v'), "| suffix:", w[len(stem):])
# b) Word generation: root + suffixes, keep only real words
root = 'play'
suffixes = ['', 's', 'ed', 'ing', 'er', 'ers', 'ful', 'ly', 'able']
generated = [root + s for s in suffixes if wordnet.synsets(root + s)]
print("Generated words:", generated)
3. Word Sense Disambiguation
we have to show 5 sentences with ambigious words and use nltk to remove ambiguity
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
import nltk
from nltk.tokenize import word_tokenize
from nltk.wsd import lesk
nltk.download('punkt')
nltk.download('punkt_tab')
samples = [('bank', "I deposited money in the bank"),
('bank', "We sat on the bank of the river"),
('bat', "He hit the ball with a bat"),
('bass', "He plays the bass guitar in a band"),
('plant', "The plant produces electricity for the city"),
('spring', "Water flows from the spring in the hill")]
for word, sent in samples:
sense = lesk(word_tokenize(sent), word)
print(word, "->", sense, ":", sense.definition() if sense else "no sense found")
4. Stemming and Lemmatization
stemming is a process of removing affixes from words to get to the root form of the word. Lemmatization is a more refined process of stemming. It takes into account the context of the word and returns the dictionary form of the word, which is also known as the lemma.
1
2
3
4
5
6
7
8
9
10
11
import nltk
from nltk.stem import PorterStemmer, LancasterStemmer, SnowballStemmer
nltk.download('punkt')
nltk.download('punkt_tab')
ws = ['running', 'flies', 'happily', 'studies', 'generously', 'connected']
porter, lancaster, snowball = PorterStemmer(), LancasterStemmer(), SnowballStemmer('english')
print(f"{'word':12}{'porter':10}{'lancaster':12}{'snowball'}")
for w in ws:
print(f"{w:12}{porter.stem(w):10}{lancaster.stem(w):12}{snowball.stem(w)}")
5. POS Tagging of words using NLTK Library
1
2
3
4
5
6
7
8
9
10
11
12
13
14
import nltk
from nltk import pos_tag
nltk.download('punkt')
nltk.download('punkt_tab')
nltk.download('averaged_perceptron_tagger')
nltk.download('averaged_perceptron_tagger_eng')
wlist = ['run', 'beautiful', 'quickly', 'university', 'she',
'and', 'under', 'eat', 'happy', 'book', 'the', 'they']
tags = dict(pos_tag(wlist))
print(tags)
w = input("Enter a word: ").lower()
print(w, "->", tags.get(w) or pos_tag([w])[0][1])
6. Morphological Analysis, N-grams, and N-gram smoothing using NLTK Library
we are given the task
- Write a Python program to
- Perform Morphological Analysis using NLTK library
- Generate n-grams using NLTK N-Grams library
- Implement N-Grams Smoothing.
and the code for that is,
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
import nltk
from nltk import ngrams, bigrams, pos_tag
from nltk.stem import PorterStemmer, WordNetLemmatizer
from nltk.corpus import wordnet
from nltk.tokenize import word_tokenize
from nltk.probability import ConditionalFreqDist, ConditionalProbDist, LaplaceProbDist
nltk.download('punkt')
nltk.download('punkt_tab')
nltk.download('wordnet')
nltk.download('averaged_perceptron_tagger_eng')
ps = PorterStemmer()
wl = WordNetLemmatizer()
# a) Morphological analysis: stem, lemma, morphy (base form), POS
for w in ['studies', 'honoring', 'better', 'cultural']:
print(w,
"| stem:", ps.stem(w),
"| lemma:", wl.lemmatize(w, 'v'),
"| morphy:", wordnet.morphy(w),
"| POS:", pos_tag([w])[0][1])
# Sentence for n-gram analysis
text = "JKKNIU stands as a dedicated center for higher education, honoring the profound cultural and literary legacy of Bangladesh's national poet."
tokens = word_tokenize(text)
words = [w.lower() for w in tokens if w.isalpha()]
# b) N-grams
for n in (1, 2, 3):
print(f"{n}-grams:", list(ngrams(words, n))[:5])
# c) Bigram smoothing: unsmoothed (MLE) vs Laplace (add-one)
corpus = "the cat sat on the mat the cat ate the fish".split()
cfd = ConditionalFreqDist(bigrams(corpus))
cpd = ConditionalProbDist(cfd, LaplaceProbDist, bins=len(set(corpus)))
print("MLE P(cat|the) =", cfd['the'].freq('cat'))
print("Laplace P(cat|the) =", cpd['the'].prob('cat'))
print("MLE P(dog|the) =", cfd['the'].freq('dog')) # 0 -> zero-probability problem
print("Laplace P(dog|the) =", cpd['the'].prob('dog')) # > 0 after smoothing
7. Audio
we are given the task,
- Write a Python program to convert audio file to text and text file to audio files.
- Using NLTK package to convert audio file to text and text file to audio files.
Text: “Jatiya Kabi Kazi Nazrul Islam University (JKKNIU) is a premier public university in Trishal, Mymensingh, established in 2006.”
- Using NLTK package to convert audio file to text and text file to audio files.
the code is,
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
import nltk
from gtts import gTTS
from pydub import AudioSegment
import speech_recognition as sr
# Input text
msg = (
"Jatiya Kabi Kazi Nazrul Islam University (JKKNIU) is a premier public university in Trishal, Mymensingh, "
"established in 2006."
)
# a) Text-to-Speech
tts = gTTS(msg, lang='en')
tts.save('jkkniu.mp3')
print("Text converted to speech: jkkniu.mp3")
# b) Convert MP3 to WAV
audio = AudioSegment.from_mp3('jkkniu.mp3')
audio.export('jkkniu.wav', format='wav')
print("MP3 converted to WAV: jkkniu.wav")
# c) Speech-to-Text
r = sr.Recognizer()
with sr.AudioFile('jkkniu.wav') as source:
audio_data = r.record(source)
text = r.recognize_google(audio_data)
print("Recognized:", text)
Complete Code for memorization is,
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
# pip install nltk gTTS SpeechRecognition pydub (pydub needs ffmpeg)
import nltk
for p in ['punkt', 'punkt_tab', 'stopwords', 'wordnet', 'omw-1.4',
'averaged_perceptron_tagger', 'averaged_perceptron_tagger_eng']:
nltk.download(p, quiet=True)
text = ("JKKNIU stands as a dedicated center for higher education, honoring "
"the profound cultural and literary legacy of Bangladesh's national poet.")
# ---------- 1. Tokenization + Stop word removal ----------
from nltk.tokenize import word_tokenize, sent_tokenize
from nltk.corpus import stopwords
tokens = word_tokenize(text)
print("Tokens:", tokens)
print("Sentences:", sent_tokenize(text))
stop = set(stopwords.words('english'))
filtered = [w for w in tokens if w.lower() not in stop and w.isalnum()]
print("After stop word removal:", filtered)
# ---------- 2. Word analysis + Word generation ----------
from nltk import FreqDist
from nltk.stem import PorterStemmer, WordNetLemmatizer
from nltk.corpus import wordnet
# a) Word analysis: frequency, length, root and affix
words = [w.lower() for w in tokens if w.isalpha()]
print("Total:", len(words), "| Unique:", len(set(words)))
print("Most common:", FreqDist(words).most_common(3))
ps, wl = PorterStemmer(), WordNetLemmatizer()
for w in ['honoring', 'dedicated', 'stands']:
stem = ps.stem(w)
print(w, "| length:", len(w), "| stem:", stem,
"| lemma:", wl.lemmatize(w, 'v'), "| suffix:", w[len(stem):])
# b) Word generation: root + suffixes, keep only real words
root = 'play'
suffixes = ['', 's', 'ed', 'ing', 'er', 'ers', 'ful', 'ly', 'able']
generated = [root + s for s in suffixes if wordnet.synsets(root + s)]
print("Generated words:", generated)
# ---------- 3. Word Sense Disambiguation (Lesk) ----------
from nltk.wsd import lesk
samples = [('bank', "I deposited money in the bank"),
('bank', "We sat on the bank of the river"),
('bat', "He hit the ball with a bat"),
('bass', "He plays the bass guitar in a band"),
('plant', "The plant produces electricity for the city"),
('spring', "Water flows from the spring in the hill")]
for word, sent in samples:
sense = lesk(word_tokenize(sent), word)
print(word, "->", sense, ":", sense.definition() if sense else "no sense found")
# ---------- 4. Stemming ----------
from nltk.stem import LancasterStemmer, SnowballStemmer
ws = ['running', 'flies', 'happily', 'studies', 'generously', 'connected']
porter, lancaster, snowball = PorterStemmer(), LancasterStemmer(), SnowballStemmer('english')
print(f"{'word':12}{'porter':10}{'lancaster':12}{'snowball'}")
for w in ws:
print(f"{w:12}{porter.stem(w):10}{lancaster.stem(w):12}{snowball.stem(w)}")
# ---------- 5. POS tagging of a word list + lookup ----------
from nltk import pos_tag
wlist = ['run', 'beautiful', 'quickly', 'university', 'she',
'and', 'under', 'eat', 'happy', 'book', 'the', 'they']
tags = dict(pos_tag(wlist))
print(tags)
w = input("Enter a word: ").lower()
print(w, "->", tags.get(w) or pos_tag([w])[0][1])
# ---------- 6. Morphology, N-grams, Smoothing ----------
from nltk import ngrams, bigrams
from nltk.probability import ConditionalFreqDist, ConditionalProbDist, LaplaceProbDist
# a) Morphological analysis: stem, lemma, morphy (base form), POS
for w in ['studies', 'honoring', 'better', 'cultural']:
print(w, "| stem:", ps.stem(w), "| lemma:", wl.lemmatize(w, 'v'),
"| morphy:", wordnet.morphy(w), "| POS:", pos_tag([w])[0][1])
# b) N-grams
for n in (1, 2, 3):
print(f"{n}-grams:", list(ngrams(words, n))[:5])
# c) Bigram smoothing: unsmoothed (MLE) vs Laplace (add-one)
corpus = "the cat sat on the mat the cat ate the fish".split()
cfd = ConditionalFreqDist(bigrams(corpus))
cpd = ConditionalProbDist(cfd, LaplaceProbDist, bins=len(set(corpus)))
print("MLE P(cat|the) =", cfd['the'].freq('cat'))
print("Laplace P(cat|the) =", cpd['the'].prob('cat'))
print("MLE P(dog|the) =", cfd['the'].freq('dog')) # 0 -> zero-probability problem
print("Laplace P(dog|the) =", cpd['the'].prob('dog')) # > 0 after smoothing
# ---------- 7. Text <-> Audio ----------
# NLTK has no audio support, so use gTTS (text->speech) and SpeechRecognition (speech->text). Needs internet.
from gtts import gTTS
from pydub import AudioSegment
import speech_recognition as sr
msg = ("Jatiya Kabi Kazi Nazrul Islam University (JKKNIU) is a premier "
"public university in Trishal, Mymensingh, established in 2006.")
# Text -> Audio
gTTS(msg, lang='en').save('jkkniu.mp3')
# Audio -> Text (recognizer reads WAV, so convert the mp3 first)
AudioSegment.from_mp3('jkkniu.mp3').export('jkkniu.wav', format='wav')
r = sr.Recognizer()
with sr.AudioFile('jkkniu.wav') as source:
print("Recognized:", r.recognize_google(r.record(source)))