การใช้การกรองสแปมแบบเบย์เชี่ยนเพื่อจำแนกอีเมลใน Python
Contents
[
Hide
]
การใช้การกรองสแปมแบบเบย์เชี่ยน
Aspose.Email ให้ฟังก์ชันการกรองอีเมลโดยใช้ Bayesian spam analyzer. มันให้ SpamAnalyzer คลาสสำหรับจุดประสงค์นี้. บทความนี้แสดงวิธีฝึกตัวกรองให้แยกแยะระหว่างสแปมและอีเมลทั่วไปโดยอิงจากฐานข้อมูลคำ
- ระบุเส้นทางโฟลเดอร์สำหรับอีเมล ham (ham_folder), อีเมลสแปม (spam_folder), อีเมลทดสอบ (test_folder), และไฟล์ฐานข้อมูล (database_file) สำหรับตัวกรองสแปม
- กำหนดฟังก์ชันช่วยเหลือ
print_resultเพื่อพิมพ์ว่าข้อความนั้นถูกจำแนกว่าเป็นสแปมหรือไม่โดยอิงจากความน่าจะเป็นสแปมที่คำนวณได้ - Create a Spam Analyzer, train it with the emails from ham_folder (not spam) and spam_folder (spam) using the ’train_filter(message, is_spam)’ method, then save the trained database with ‘save_database(file_path)’.
- Create a Spam Analyzer, restore the trained database with ’load_database(file_path)’, load the .eml files from ’test_folder’, analyze each with ’test(message)’ to get the spam probability, and print the email subject and classification using ‘print_result’.
import os
from aspose.email import MailMessage
from aspose.email.antispam import SpamAnalyzer
ham_folder = "hamFolder"
spam_folder = "spamFolder"
test_folder = "testFolder"
database_file = "SpamFilterDatabase.txt"
def print_result(probability):
if probability >= 0.5:
print("The message is classified as spam.")
else:
print("The message is classified as not spam.")
print("Spam Probability: " + str(probability))
print()
def train_from_directory(analyzer, folder, is_spam):
for file in os.listdir(folder):
if file.endswith(".eml"):
analyzer.train_filter(MailMessage.load(os.path.join(folder, file)), is_spam)
def teach_and_create_database(ham_folder, spam_folder, database_file):
analyzer = SpamAnalyzer()
train_from_directory(analyzer, ham_folder, False)
train_from_directory(analyzer, spam_folder, True)
analyzer.save_database(database_file)
teach_and_create_database(ham_folder, spam_folder, database_file)
test_files = [f for f in os.listdir(test_folder) if f.endswith(".eml")]
analyzer = SpamAnalyzer()
analyzer.load_database(database_file)
for file in test_files:
file_path = os.path.join(test_folder, file)
msg = MailMessage.load(file_path)
print(msg.subject)
probability = analyzer.test(msg)
print_result(probability)