FMM.py

(开头部分) 1KB

这里只显示每个文件的开头 60 行。登录后可以解锁完整代码。

#!/usr/bin/python3
# -*- coding: UTF-8 -*-
""" 
@author: Colin
@file: FMM.py
@time: 2018/08/18
""" 
#使用正向最大匹配算法实现中文分词

dic = []

def init():
    """
    读取文件
    获取中文字典
    :return:
    """
    with open("D:/MillionDict.txt","r",encoding="utf-8") as words_dic:
        for word in words_dic:
             dic.append(word.replace("\t","").replace("\n",""))

def if_contain(words):
    """
    判断当前词是否在词典中
    :param words:
    :return:
    """
    flag=False
    for word in dic:
        if word == words:
            flag = True
            break
    return flag

def bmm(sentence):
     """
    正向最大匹配算法主要实现
    从前往后切割字符串,直到切割出的字符串与字典中的词匹配
    :param sentence:
    :return:
    """
     init()
     result = ''
     words = []

     while len(sentence) > 0:
         except_flag = False
         for i in range(0,len(sentence)-1 ,1):
             temp = sentence[i:]
             flag = if_contain(temp)
             if flag:
                 words.append(temp)
                 sentence = sentence[:i]
                 except_flag = True
                 break
         if not except_flag:
             """
           判断当前字符串是否在字典中不存在,该字符串从头到尾在字典中都不存在,则直接放入最后结果中
           """
             words.append(sentence)
后面还有 5 行代码,购买后查看完整代码

24 小时内免费解锁 3 个项目,之后 1 积分/个。 规则说明

AI 解读

登录后可用,每次 10 积分,解读结果公开显示在下面。

还没有人解读过这个文件。