pyspark SparseVector 词向量

转载

luoganttcc 2023-01-13 00:21:30 博主文章分类：spark

文章标签 spark 文章分类 运维

from pyspark.mllib.linalg import SparseVector
from collections import Counter

from pyspark import SparkContext

if __name__ == "__main__":

    sc = SparkContext('local', 'term_doc')
    corpus = sc.parallelize([
    "It is the east, and Juliet is the sun.",
    "A dish fit for the gods.",
    "Brevity is the soul of wit."])

    tokens = corpus.map(lambda raw_text: raw_text.split()).cache()   
    local_vocab_map = tokens.flatMap(lambda token: token).distinct().zipWithIndex().collectAsMap()

    vocab_map = sc.broadcast(local_vocab_map)
    vocab_size = sc.broadcast(len(local_vocab_map))

    term_document_matrix = tokens \
                         .map(Counter) \
                         .map(lambda counts: {vocab_map.value[token]: float(counts[token]) for token in counts}) \
                         .map(lambda index_counts: SparseVector(vocab_size.value, index_counts))

    for doc in term_document_matrix.collect():
        print( doc)

(16,[0,1,2,3,4,5,6],[1.0,2.0,2.0,1.0,1.0,1.0,1.0])
(16,[2,7,8,9,10,11],[1.0,1.0,1.0,1.0,1.0,1.0])
(16,[1,2,12,13,14,15],[1.0,1.0,1.0,1.0,1.0,1.0])

上一篇：pyspark parallelize

下一篇：pyspark filter

提问和评论都可以，用心的回复会被更多人看到评论

发布评论

相关文章

官方博客	全部文章	热门标签	班级博客
了解我们	网站地图	意见反馈

鸿蒙开发者社区	51CTO学堂
51CTO	软考资讯

pyspark SparseVector 词向量

pyspark SparseVector 词向量

51CTO博客