在 for 循环中并行化函数

在您的函数中，您可以决定通过将文本拆分为子部分来并行化，将标记化应用于子部分，然后连接结果。沿线的东西：text0 = text[:len(text)/2]text1 = text[len(text)/2:]然后将您的处理应用于这两部分，使用：# here, I suppose that clean_preprocess is the sequential version, # and we manage the pool outside of itwith Pool(2) as p:  words0, words1 = pool.map(clean_preprocess, [text0, text1])words = words1 + words2# or continue with words0 words1 to save the cost of joining the lists然而，你的函数似乎受内存限制，所以它不会有一个可怕的加速（通常，因子 2 是我们现在在标准计算机上希望的最大值），请参阅例如，如果程序是内存，并行化对性能有多大帮助-边界？或者术语“CPU 绑定”和“I/O 绑定”是什么意思？因此，您可以尝试将文本分成 2 个以上的部分，但可能不会变得更快。您甚至可能会得到令人失望的性能，因为拆分文本可能比处理文本更昂贵。

在 for 循环中并行化函数

1回答