Python中国的学习方式处理问题

转载

mob60475705c8db 2015-09-23 11:42:00

文章标签 ico python 无法识别 windows系统默认编码 文章分类 Java 后端开发

a = '你们' 至 str 物

a = u'你们' 至 unicode 物

>>> print 'u' + '你们'

>>> u欢

输出乱码

>>> print 'u' + u'你'

>>> u你

正常

>>> print 'u你'

>>> u浣

输出乱码

>>> print 'u你' + 'u'

>>> u浣爑

输出乱码

>>> print u'u你' + 'u'

>>> u你u

正常

>>> print u'u你' + '你'

出现错误 UnicodeDecodeError: 'ascii' codec can't decode byte 0xe4 in position 0: ordinal not in range(128)

分析：'你'在内存中为 0xe4。而python默认的编码方案是ascii，ascii无法识别0xe4

>>> print u'u你' + u'你'

>>> u你你

正常

>>> print 'u你' + u'你'

出现错误 UnicodeDecodeError: 'ascii' codec can't decode byte 0xe4 in position 1: ordinal not in range(128)

>>> print 'u你'.decode('utf-8') + u'你'

>>> u你你

正常

10.

而在处理由系统採集的含有中文的路径时，使用string.decode('utf-8')就不一定行了，由于中文简体的windows系统默认编码为gb2312，繁体中文版会採用Big5码

实验步骤例如以下：

file_from = sys.argv[1] 为由系统採集的包括中文的路径

file_to = file_from[:file_from.rfind('\\')+1].decode('utf-8') + u'你_' + file_from[file_from.rfind('\\')+1:].decode('utf-8')

print file_to

将出现错误：UnicodeDecodeError: 'utf8' codec can't decode byte 0xbb in position 24: invalid start byte

应该使用：decode('gb2312')

file_to = file_from[:file_from.rfind('\\')+1].decode('gb2312') + u'你_' + file_from[file_from.rfind('\\')+1:].decode('gb2312')

print file_to 正常

11.

而假设file_from是由你自己写入的包括中文的路径，如file_from = ‘c:\你.txt’

那么就应该用decode('utf-8')

能够參考上面的第7点和第9点

本文章为转载内容，我们尊重原作者对文章享有的著作权。如有内容错误或侵权问题，欢迎原作者联系我们进行内容更正或删除文章。

上一篇：按钮点击效果(水波纹)

下一篇：百度地图如何显示缩放控件以及缩放控件的位置？

提问和评论都可以，用心的回复会被更多人看到评论

发布评论

相关文章

官方博客	全部文章	热门标签	班级博客
了解我们	网站地图	意见反馈

鸿蒙开发者社区	51CTO学堂
51CTO	软考资讯

Python中国的学习方式处理问题

Python中国的学习方式处理问题

51CTO博客