Some statistics on the corpus resulting from a preliminary exploration using Voyant.
An example of what the data looks like. In the process of cleaning the data, I removed the blurb
between the actual content of speech and the title of the speech. This blurb was generally an introduction
by someone else of Mark Twain, and it didn't seem relevant to understanding the topics discussed in the
speech.
This picture shows four abstract topics extracted from Mark Twain's speeches. The number
four is arbitrary because you can specify as many topics as you want, although the topics may become
less coherent if you increase the number too much. As we see here, abstract Topic 1 can be intepreted
as content related to the speech itself. For example, when Mark Twain acknowledges the one introducing
him. Abstract Topic 2, we may be able to intepret it as his pride of being an American. Abstract Topic 3 could
be interpreted as Mark Twain addressing the current status of America at the time. Topic 4 likely falls under
the umbrella of business. This makes sense as, in addition to a many other things, Mark Twain was a businessman.