java 读取doc文件内容 java读取word文档

转载

deanyuancn 2023-06-21 16:50:04

文章标签 JAVA如何读取docx文件指定内容 java 用pdf格式读取word apache jar JAVA 文章分类 Java 后端开发

本篇文章主要通过实例代码介绍了JAVA读取PDF、WORD文档，需要的朋友可以参考下

读取PDF文件jar引用

org.apache.pdfbox
pdfbox
1.8.13

读取WORD文件jar引用

org.apache.poi
poi-scratchpad
3.16-beta1
org.apache.poi
poi
3.16-beta1

读取WORD文件方法/**

*
* @Title: getTextFromWord
* @Description: 读取word
* @param filePath
* 文件路径
* @return: String 读出的Word的内容
*/
public static String getTextFromWord(String filePath) {
String result = null;
File file = new File(filePath);
FileInputStream fis = null;
try {
fis = new FileInputStream(file);
@SuppressWarnings("resource")
WordExtractor wordExtractor = new WordExtractor(fis);
result = wordExtractor.getText();
} catch (FileNotFoundException e) {
e.printStackTrace();
} catch (IOException e) {
e.printStackTrace();
} finally {
if (fis != null) {
try {
fis.close();
} catch (IOException e) {
e.printStackTrace();
}
}
}
return result;
}

读取PDF文件方法

/**
*
* @Title: getTextFromPdf
* @Description: 读取pdf文件内容
* @param filePath
* @return: 读出的pdf的内容
*/
public static String getTextFromPdf(String filePath) {
String result = null;
FileInputStream is = null;
PDDocument document = null;
try {
is = new FileInputStream(filePath);
PDFParser parser = new PDFParser(is);
parser.parse();
document = parser.getPDDocument();
PDFTextStripper stripper = new PDFTextStripper();
result = stripper.getText(document);
} catch (FileNotFoundException e) {
e.printStackTrace();
} catch (IOException e) {
e.printStackTrace();
} finally {
if (is != null) {
try {
is.close();
} catch (IOException e) {
e.printStackTrace();
}
}
if (document != null) {
try {
document.close();
} catch (IOException e) {
e.printStackTrace();
}
}
}
return result;
}

本文章为转载内容，我们尊重原作者对文章享有的著作权。如有内容错误或侵权问题，欢迎原作者联系我们进行内容更正或删除文章。