class

Cadmium::Classifier::Bayes

Inherits MessagePack::Serializable < YAML::Serializable < JSON::Serializable < Reference < Object

This is a native-bayes classifier which used Laplace Smoothing. It can be trained to categorize sentences based on the words in that sentence.

Example:

classifier = Cadmium::Classifier::Bayes.new

# Train some angry examples
classifier.train("omg I can't believe you would do that to me", "angry")
classifier.train("I hate you so much!", "angry")
classifier.train("Just go. I don't need this.", "angry")
classifier.train("You're so full of shit!", "angry")

# Some happy ones
classifier.train("omg you're the best!", "happy")
classifier.train("I can't believe how happy you make me", "happy")
classifier.train("I love you so damn much!", "happy")
classifier.train("You're the best!", "happy")

# And some indifferent ones
classifier.train("Idk, what do you think?", "indifferent")
classifier.train("yeah that's ok", "indifferent")
classifier.train("cool", "indifferent")
classifier.train("I guess we could do that", "indifferent")

# Now let's test it on a sentence
classifier.classify("You shit head!")
# => {"angry" => 85.5, "happy" => 10.2, "indifferent" => 4.3}

# Or just get the top category
classifier.classify_category("You're the best :)")
# => "happy"

Constants

DEFAULT_TOKENIZER = Cadmium::Tokenizer::Word.new

Constructors

new(tokenizer = nil)
Source
new(*, __pull_for_json_serializable pull : JSON::PullParser)
Source
new(*, __context_for_yaml_serializable ctx : YAML::ParseContext, __node_for_yaml_serializable node : YAML::Nodes::Node)
Source

Instance methods

categories

Category names

Source
classify(text : String)

Determines what category the text belongs to. Returns a Hash with all categories and their probabilities.

Source
classify_category(text : String) : String

Convenience method that returns just the top category name instead of all probabilities. Use this when you only need the most likely category.

Example:

classifier = Cadmium::Classifier::Bayes.new
classifier.train("I love this!", "positive")
classifier.train("This is terrible", "negative")
classifier.classify_category("This is amazing!") # => "positive"
Source
doc_count

Document frequency table for each of our categories.

Source
frequency_table(tokens)

Build a frequency hash map where

  • the keys are the entries in tokens
  • the values are the frequency of each entry in tokens
Source
initialize_category(name)

Intializes each of our data structure entities for this new category and returns self.

Source
token_probability(token, category)

Calculate the probaility that a token belongs to a category.

Source
tokenizer
Source
tokenizer=(tokenizer : Cadmium::Tokenizer::Base)
Source
total_documents

Number of documents we have learned from.

Source
train(text, category)

Train our native-bayes classifier by telling it what category the train text corresponds to.

Source
vocabulary

The words to learn from. Using Set for O(1) lookups instead of Array's O(n).

Source
vocabulary_size

The total number of words in the vocabulary

Source
word_count

For each category, how many total words were mapped to it.

Source
word_frequency_count

Word frequency table for each category.

Source