Sử dụng lọc thư rác Bayesian để phân loại email trong Python

Sử dụng Lọc Spam Bayesian

Aspose.Email cung cấp chức năng lọc email bằng bộ phân tích spam Bayesian. Nó cung cấp SpamAnalyzer lớp cho mục đích này. Bài viết này trình bày cách huấn luyện bộ lọc để phân biệt giữa spam và email thông thường dựa trên cơ sở dữ liệu từ.

  1. Chỉ định đường dẫn thư mục cho các email ham (ham_folder), email spam (spam_folder), email kiểm tra (test_folder), và tệp cơ sở dữ liệu (database_file) cho bộ lọc spam.
  2. Định nghĩa hàm trợ giúp print_result để in ra thông báo liệu một tin nhắn có được phân loại là spam hay không dựa trên xác suất spam đã tính.
  3. Tạo một Spam Analyzer, đào tạo nó với các email từ ham_folder (không phải spam) và spam_folder (spam) bằng phương thức ’train_filter(message, is_spam)’, sau đó lưu cơ sở dữ liệu đã huấn luyện bằng ‘save_database(file_path)’.
  4. Tạo một Spam Analyzer, khôi phục cơ sở dữ liệu đã huấn luyện bằng ’load_database(file_path)’, tải các tệp .eml từ ’test_folder’, phân tích từng tệp bằng ’test(message)’ để nhận xác suất spam, và in tiêu đề email cùng phân loại bằng ‘print_result’.
import os

from aspose.email import MailMessage
from aspose.email.antispam import SpamAnalyzer

ham_folder = "hamFolder"
spam_folder = "spamFolder"
test_folder = "testFolder"
database_file = "SpamFilterDatabase.txt"

def print_result(probability):
    if probability >= 0.5:
        print("The message is classified as spam.")
    else:
        print("The message is classified as not spam.")
    print("Spam Probability: " + str(probability))
    print()

def train_from_directory(analyzer, folder, is_spam):
    for file in os.listdir(folder):
        if file.endswith(".eml"):
            analyzer.train_filter(MailMessage.load(os.path.join(folder, file)), is_spam)

def teach_and_create_database(ham_folder, spam_folder, database_file):
    analyzer = SpamAnalyzer()
    train_from_directory(analyzer, ham_folder, False)
    train_from_directory(analyzer, spam_folder, True)
    analyzer.save_database(database_file)

teach_and_create_database(ham_folder, spam_folder, database_file)

test_files = [f for f in os.listdir(test_folder) if f.endswith(".eml")]
analyzer = SpamAnalyzer()
analyzer.load_database(database_file)

for file in test_files:
    file_path = os.path.join(test_folder, file)
    msg = MailMessage.load(file_path)
    print(msg.subject)
    probability = analyzer.test(msg)
    print_result(probability)