FMM.py
(开头部分) 1KB这里只显示每个文件的开头 60 行。登录后可以解锁完整代码。
#!/usr/bin/python3
# -*- coding: UTF-8 -*-
"""
@author: Colin
@file: FMM.py
@time: 2018/08/18
"""
#使用正向最大匹配算法实现中文分词
dic = []
def init():
"""
读取文件
获取中文字典
:return:
"""
with open("D:/MillionDict.txt","r",encoding="utf-8") as words_dic:
for word in words_dic:
dic.append(word.replace("\t","").replace("\n",""))
def if_contain(words):
"""
判断当前词是否在词典中
:param words:
:return:
"""
flag=False
for word in dic:
if word == words:
flag = True
break
return flag
def bmm(sentence):
"""
正向最大匹配算法主要实现
从前往后切割字符串,直到切割出的字符串与字典中的词匹配
:param sentence:
:return:
"""
init()
result = ''
words = []
while len(sentence) > 0:
except_flag = False
for i in range(0,len(sentence)-1 ,1):
temp = sentence[i:]
flag = if_contain(temp)
if flag:
words.append(temp)
sentence = sentence[:i]
except_flag = True
break
if not except_flag:
"""
判断当前字符串是否在字典中不存在,该字符串从头到尾在字典中都不存在,则直接放入最后结果中
"""
words.append(sentence)
后面还有 5 行代码,购买后查看完整代码
24 小时内免费解锁 3 个项目,之后 1 积分/个。 规则说明
AI 解读
登录后可用,每次 10 积分,解读结果公开显示在下面。
还没有人解读过这个文件。
