Class 10 — Monday, September 28
Section 1 runs on the projector — keyboards down, prediction in the Journal first. Section 2 is typed on the projector by the room, one line each. From section 3 you are in pairs, in any project you have open.
- Three sets, three orders — predict first
- Count the words, together
- The five commonest
- The words two texts share
- One word changed
- Which lines is it on?
- Words by how often
- The same letters
- Pairs of words
retainAllby hand- The longest words
- One loop instead of two
Sections 1 to 4 are the class. Sections 5 to 12 go past anything you have been asked for — take them in order if you get there.
The two texts, for sections 2 to 6. Save them next to your project, not inside src:
moby.txt— the opening of Moby-Dick, abridged.austen.txt— the opening of Pride and Prejudice, abridged.
1. Three sets, three orders
Keyboards down. Prediction before it runs.
Seven words, three sets. The third set is told how to compare, by a comparator of the kind you wrote for A4:
class ByLength implements Comparator<String> {
@Override
public int compare(String a, String b) {
return Integer.compare(a.length(), b.length());
}
}List<String> words = List.of("sea", "ship", "ocean", "purse", "hat", "coffin", "sword");
Set<String> hashed = new HashSet<>(words);
Set<String> sorted = new TreeSet<>(words);
Set<String> byLength = new TreeSet<>(new ByLength());
byLength.addAll(words);
System.out.println(hashed);
System.out.println(sorted);
System.out.println(byLength);
System.out.println(hashed.size() + " " + sorted.size() + " " + byLength.size());Write your answer in the Journal before it runs. Four lines. A guess counts.
What are the three lists, and what are the three sizes?
[sword, ocean, ship, hat, coffin, sea, purse]
[coffin, hat, ocean, purse, sea, ship, sword]
[sea, ship, ocean, coffin]
7 7 4
Nobody could have predicted the first line, and that is the point. A HashSet promises no order at all. What you see is where the hash codes sent each word, and the next JDK may print it differently.
The second line is alphabetical, because a TreeSet asks compareTo, and String has one.
The third set threw three words away. It was told to compare by length, so hat and sea are the same word to it, and so are ship, purse and sword. Same seven words. Nobody changed them. The set decides what counts as the same, and you told this one how.
One word changed. Make the first set a LinkedHashSet instead, add "sea" to it a second time, and predict the line and the size before it runs.
[sea, ship, ocean, purse, hat, coffin, sword] 7
The order you added them in, and the second sea changed nothing. A LinkedHashSet is a HashSet that also remembers the order of arrival — same rule about duplicates, one more promise about order.
2. Count the words, together
Keyboards down. You dictate, one line each; it is typed on the projector.
The loop is written. It hands you one word at a time, lower-cased, no punctuation:
for (String line : Files.readAllLines(Path.of("moby.txt"))) {
StringTokenizer words = new StringTokenizer(line, " ,.;:?!\"()");
while (words.hasMoreTokens()) {
String w = words.nextToken().toLowerCase();
// w is one word
}
}A StringTokenizer cuts a line wherever it meets one of the characters you list, and hands you the pieces in between. toLowerCase makes Call and call the same word. You need java.util.StringTokenizer, java.nio.file.Files and java.nio.file.Path.
What the room writes: a Map<String, Integer> before the loop, one line inside it that adds one to this word’s count, and a print of the map’s size() after it.
Stop before it runs. The line most rooms write first is
counts.put(w, counts.get(w) + 1);Does it print a number, or does it crash?
Exception in thread "main" java.lang.NullPointerException:
Cannot invoke "java.lang.Integer.intValue()"
because the return value of "java.util.Map.get(Object)" is null
The first word is not in the map yet. get says so by returning null, and null + 1 has to turn null into an int first. That is the Integer against int question from Wednesday, thrown at you by a map.
getOrDefault(w, 0) is get with an answer for the missing case:
counts.put(w, counts.getOrDefault(w, 0) + 1);134
That is the number of distinct words: the map holds each word once, whatever put was called. Two hundred and one words went past, and a word already in the map had its count replaced, not a second entry added.
3. The five commonest
Keyboards up. Pairs. Type section 2’s program into any project, with moby.txt beside it.
A map has no order. To print the five commonest words you have to make one: copy the keys into a list, sort the list with a comparator that looks the counts up, and print the first five.
List<String> ranked = new ArrayList<>(counts.keySet());
ranked.sort(new ByCount(counts));ByCount is a Comparator<String> that needs to see the map, so give it the map in its constructor and keep it in a field. Most frequent first; two words with the same count go alphabetically.
Done when you get
the 10
i 9
and 7
a 5
it 5
a and it both appear five times, so the count comparison says they tie, and the alphabetical rule is what puts a first. Without that rule the sort would leave tied words in whatever order the map handed them over, which can differ from one program to the next — two pairs could both be right and print different lines. The rule is what lets your output be checked against this one.
5. One word changed
Change HashMap to TreeMap in section 2 and print the first six entries.
a=5 about=2 account=1 ago=1 all=1 almost=1
Same counts, and now the keys come out sorted, because a TreeMap keeps its keys the way a TreeSet keeps its elements. Every line of the program that touches the map is unchanged. The interface is the promise; the class you picked is how it keeps it.
6. Which lines is it on?
This is the one to get to.
A map’s value can be anything, including a list. Make a Map<String, List<Integer>> from word to the line numbers it appears on, and print the lists for whenever, sea and whale.
[4, 5, 6, 8] [11] null
The first time a word turns up, there is no list to add to yet. That is the section 2 crash waiting to happen again, and this time getOrDefault is not enough on its own, because the default has to be a list the map keeps. Work out the two lines that put a new list in and add to it.
whale prints null, and the whole paragraph is about one. Say why.
7. Words by how often
Turn the count map inside out: a map from a count to the set of words that appear exactly that many times. A TreeMap<Integer, TreeSet<String>> keeps the counts in order and each set alphabetical. The first time you see a count there is no set yet — section 6 again.
Done when you can print how many words appear exactly once, and every word that appears five times or more:
109
{5=[a, it, me, to], 7=[and], 9=[i], 10=[the]}
tailMap(5) is the part of a TreeMap from a key onwards. Read what it returns before you use it.
8. The same letters
Two words are anagrams when they are the same letters in a different order. Sort a word’s letters and you have a key that both words share:
char[] letters = w.toCharArray();
Arrays.sort(letters);
String key = new String(letters);Map that key to the set of words that produce it, over both texts together, then print every set with more than one word in it.
Done when you get
[how, who]
[no, on]
in either order. Moby-Dick alone has only the second pair.
9. Pairs of words
Count pairs of consecutive words rather than words: call me, me ishmael, and so on through the whole text. The key is the two words with a space between. You need the previous word, and it has to survive from one line of the file to the next.
Done when you can say how many distinct pairs there are and print the commonest ones:
194
[find myself, i find, in my, is a, it is, whenever i]
Six pairs tie at two apiece. Predict before you run it: did you expect any pair to appear more than twice in a paragraph this size?
10. retainAll by hand
Section 4 called retainAll. Write it yourself: a method that takes a set of Moby’s words and a set of Austen’s, and removes from the first every word that is not in the second. You are removing while walking, so the for loop from Wednesday cannot do it and an Iterator can.
static void keepOnly(Set<String> words, Set<String> keep)Done when your method leaves 34 words in the set and words.equals(shared) prints true against the set retainAll produced in section 4. Two sets are equals when they hold the same elements, whatever kind of set each is.
11. The longest words
A map from a length to the set of words of that length, in a TreeMap, and then its last entry.
Done when you get
13=[involuntarily, philosophical]
lastEntry() is the TreeMap method. Then answer this in a comment: why is a TreeMap the right choice here and a HashMap the wrong one, when both would hold exactly the same pairs?
12. One loop instead of two
Section 2 reads the file a line at a time and then cuts each line into words: two loops. Only section 6 needs to know which line a word is on, so everywhere else one loop should do.
Write section 2 again with one loop. Files.readString hands you the whole file as a single String. Then look at what now sits between the last word of one line and the first word of the next, and at the list of characters you gave your StringTokenizer.
Done when it counts 201 words, 134 of them different — the same as section 2.
Print each word between square brackets, "[" + w + "]", and look at the ones that come out strangely. Something is stuck to them that you cannot see.