使用PDFBox處理PDF文件

阿新 • • 發佈：2019-01-10

專案需要在原有的PDF檔案中插入圖片、文字，並將最終的PDF檔案轉換為圖片，在網上找了很多Demo，現在開源可以解析處理PDF檔案的第三方外掛比較多，eg：IText、PDFBox等，現在就PDFBox解析處理PDF檔案總結如下：

【PDFBox簡介】

自從Adobe公司1993年第一次釋出公共PDF參考以來，支援各種語言和平臺的PDF工具和類庫就如雨後春筍般湧現。然而，Java應用開發中Adobe技術的支援相對滯後了。這是個奇怪的現象，因為PDF文件是企業資訊系統儲存和交換資訊的大勢所趨，而Java技術特別適合這種應用。然而，Java開發人員似乎直到最近才獲得成熟可用的PDF支援。 PDFBox（一個BSD許可下的原始碼開放專案）是一個為開發人員讀取和建立PDF文件而準備的純Java類庫。它提供如下特性：提取文字，包括Unicode字元。和Jakarta Lucene等文字搜尋引擎的整合過程十分簡單。加密/解密PDF文件。從PDF和XFDF格式中匯入或匯出表單資料。向已有PDF文件中追加內容。將一個PDF文件切分為多個文件。覆蓋PDF文件。 PS：http://baike.baidu.com/link?url=TsYWHJtTPMhlf0UvKzPOk-j3f9KzF7morIa4CqoZ0s4yIDCLB3z8nLVgLHVz-AO4dE6S7ls_3_yuvXP03nLSiq

【PDFBox下載】

最常見的一種PDF文字抽取工具就是PDFBox了，訪問網址http://pdfbox.apache.org/download.cgi，進入如下圖所示的下載介面。讀者可以在該網頁下載其最新的版本。本書採用的是pdfbox-1.8.8版本。PDFBox是一個開源的Java PDF庫，這個庫允許你訪問PDF檔案的各項資訊。在接下來的例子中，將演示如何使用PDFBox提供的API操作PDF檔案。

【將剛下載的7個jar包引入到工程當中】

pdfbox-1.8.8-src.zip為pdfbox原始碼，裡面有很對的例子，在pdfbox-1.8.8\examples目錄下存在

【以下為Demo正式開始】

1、建立PDF檔案

 1     public void createHelloPDF() {
 2         PDDocument doc = null;
 3         PDPage page = null;
 4 
 5         try {
 6             doc = new PDDocument();
 7             page = new PDPage();
 8             doc.addPage(page);
 9             PDFont font = PDType1Font.HELVETICA_BOLD;
 
10             PDPageContentStream content = new PDPageContentStream(doc, page);
11             content.beginText();
12             content.setFont(font, 12);
13             content.moveTextPositionByAmount(100, 700);
14             content.drawString("hello");
15 
16             content.endText();
17             content.close();
18             doc.save("F:\\java56班\\eclipse-SDK-4.2-win32\\pdfwithText.pdf");
19             doc.close();
20         } catch (Exception e) {
21             System.out.println(e);
22         }
23     }

2、讀取PDF檔案：

 1     public void readPDF() {
 2         PDDocument helloDocument = null;
 3         try {
 4             helloDocument = PDDocument.load(new File(
 5                     "F:\\java56班\\eclipse-SDK-4.2-win32\\pdfwithText.pdf"));
 6             PDFTextStripper textStripper = new PDFTextStripper("GBK");
 7             System.out.println(textStripper.getText(helloDocument));
 8 
 9             helloDocument.close();
10         } catch (IOException e) {
11             // TODO Auto-generated catch block
12             e.printStackTrace();
13         }
14     }

3、修改PDF檔案（處理中文亂碼，我可以搞定的）：

 1  /**
 2      * Locate a string in a PDF and replace it with a new string.
 3      *
 4      * @param inputFile The PDF to open.
 5      * @param outputFile The PDF to write to.
 6      * @param strToFind The string to find in the PDF document.
 7      * @param message The message to write in the file.
 8      *
 9      * @throws IOException If there is an error writing the data.
10      * @throws COSVisitorException If there is an error writing the PDF.
11      */
12     public void doIt( String inputFile, String outputFile, String strToFind, String message)
13         throws IOException, COSVisitorException
14     {
15         // the document
16         PDDocument doc = null;
17         try
18         {
19             doc = PDDocument.load( inputFile );
20 //            PDFTextStripper stripper=new PDFTextStripper("ISO-8859-1");
21             List pages = doc.getDocumentCatalog().getAllPages();
22             for( int i=0; i<pages.size(); i++ )
23             {
24                 PDPage page = (PDPage)pages.get( i );
25                 PDStream contents = page.getContents();
26                 PDFStreamParser parser = new PDFStreamParser(contents.getStream() );
27                 parser.parse();
28                 List tokens = parser.getTokens();
29                 for( int j=0; j<tokens.size(); j++ )
30                 {
31                     Object next = tokens.get( j );
32                     if( next instanceof PDFOperator )
33                     {
34                         PDFOperator op = (PDFOperator)next;
35                         //Tj and TJ are the two operators that display
36                         //strings in a PDF
37                         if( op.getOperation().equals( "Tj" ) )
38                         {
39                             //Tj takes one operator and that is the string
40                             //to display so lets update that operator
41                             COSString previous = (COSString)tokens.get( j-1 );
42                             String string = previous.getString();
43                             string = string.replaceFirst( strToFind, message );
44                             System.out.println(string);
45                             System.out.println(string.getBytes("GBK"));
46                             previous.reset();
47                             previous.append( string.getBytes("GBK") );
48                         }
49                         else if( op.getOperation().equals( "TJ" ) )
50                         {
51                             COSArray previous = (COSArray)tokens.get( j-1 );
52                             for( int k=0; k<previous.size(); k++ )
53                             {
54                                 Object arrElement = previous.getObject( k );
55                                 if( arrElement instanceof COSString )
56                                 {
57                                     COSString cosString = (COSString)arrElement;
58                                     String string = cosString.getString();
59                                     string = string.replaceFirst( strToFind, message );
60                                     cosString.reset();
61                                     cosString.append( string.getBytes("GBK") );
62                                 }
63                             }
64                         }
65                     }
66                 }
67                 //now that the tokens are updated we will replace the
68                 //page content stream.
69                 PDStream updatedStream = new PDStream(doc);
70                 OutputStream out = updatedStream.createOutputStream();
71                 ContentStreamWriter tokenWriter = new ContentStreamWriter(out);
72                 tokenWriter.writeTokens( tokens );
73                 page.setContents( updatedStream );
74             }
75             doc.save( outputFile );
76         }
77         finally
78         {
79             if( doc != null )
80             {
81                 doc.close();
82             }
83         }
84     }

4、在PDF中加入圖片：

 1 /**
 2      * Add an image to an existing PDF document.
 3      *
 4      * @param inputFile The input PDF to add the image to.
 5      * @param image The filename of the image to put in the PDF.
 6      * @param outputFile The file to write to the pdf to.
 7      *
 8      * @throws IOException If there is an error writing the data.
 9      * @throws COSVisitorException If there is an error writing the PDF.
10      */
11     public void createPDFFromImage( String inputFile, String image, String outputFile ) 
12         throws IOException, COSVisitorException
13     {
14         // the document
15         PDDocument doc = null;
16         try
17         {
18             doc = PDDocument.load( inputFile );
19 
20             //we will add the image to the first page.
21             PDPage page = (PDPage)doc.getDocumentCatalog().getAllPages().get( 0 );
22 
23             PDXObjectImage ximage = null;
24             if( image.toLowerCase().endsWith( ".jpg" ) )
25             {
26                 ximage = new PDJpeg(doc, new FileInputStream( image ) );
27             }
28             else if (image.toLowerCase().endsWith(".tif") || image.toLowerCase().endsWith(".tiff"))
29             {
30                 ximage = new PDCcitt(doc, new RandomAccessFile(new File(image),"r"));
31             }
32             else
33             {
34                 BufferedImage awtImage = ImageIO.read( new File( image ) );
35                 ximage = new PDPixelMap(doc, awtImage);
36             }
37             PDPageContentStream contentStream = new PDPageContentStream(doc, page, true, true);
38 
39             //contentStream.drawImage(ximage, 20, 20 );
40             // better method inspired by http://stackoverflow.com/a/22318681/535646
41             float scale = 0.5f; // reduce this value if the image is too large
42             System.out.println(ximage.getHeight());
43             System.out.println(ximage.getWidth());
44 //            ximage.setHeight(ximage.getHeight()/5);
45 //            ximage.setWidth(ximage.getWidth()/5);
46             contentStream.drawXObject(ximage, 20, 200, ximage.getWidth()*scale, ximage.getHeight()*scale);
47 
48             contentStream.close();
49             doc.save( outputFile );
50         }
51         finally
52         {
53             if( doc != null )
54             {
55                 doc.close();
56             }
57         }
58     }

5、PDF檔案轉換為圖片：

 1 public void toImage() {
 2     try {
 3         PDDocument doc = PDDocument
 4                 .load("F:\\java56班\\eclipse-SDK-4.2-win32\\pdfwithText.pdf");
 5         int pageCount = doc.getPageCount();
 6         System.out.println(pageCount);
 7         List pages = doc.getDocumentCatalog().getAllPages();
 8         for (int i = 0; i < pages.size(); i++) {
 9             PDPage page = (PDPage) pages.get(i);
10             BufferedImage image = page.convertToImage();
11             Iterator iter = ImageIO.getImageWritersBySuffix("jpg");
12             ImageWriter writer = (ImageWriter) iter.next();
13             File outFile = new File("F:\\java56班\\eclipse-SDK-4.2-win32\\"
14                     + i + ".jpg");
15             FileOutputStream out = new FileOutputStream(outFile);
16             ImageOutputStream outImage = ImageIO
17                     .createImageOutputStream(out);
18             writer.setOutput(outImage);
19             writer.write(new IIOImage(image, null, null));
20         }
21         doc.close();
22         System.out.println("over");
23     } catch (FileNotFoundException e) {
24         // TODO Auto-generated catch block
25         e.printStackTrace();
26     } catch (IOException e) {
27         // TODO Auto-generated catch block
28         e.printStackTrace();
29     }
30 }

6、圖片轉換為PDF檔案（支援多張圖片轉換為PDF檔案）：

 1 /**
 2      * create the second sample document from the PDF file format specification.
 3      * 
 4      * @param file
 5      *            The file to write the PDF to.
 6      * @param image
 7      *            The filename of the image to put in the PDF.
 8      * 
 9      * @throws IOException
10      *             If there is an error writing the data.
11      * @throws COSVisitorException
12      *             If there is an error writing the PDF.
13      */
14     public void createPDFFromImage(String file, String image)throws IOException, COSVisitorException {
15         // 多張圖片轉換為PDF檔案
16         PDDocument doc = null;
17         doc = new PDDocument();
18         PDPage page = null;
19         PDXObjectImage ximage = null;
20         PDPageContentStream contentStream = null;
21 
22         File files = new File(image);
23         String[] a = files.list();
24         for (String string : a) {
25             if (string.toLowerCase().endsWith(".jpg")) {
26                 String temp = image + "\\" + string;
27                 ximage = new PDJpeg(doc, new FileInputStream(temp));
28                 page = new PDPage();
29                 doc.addPage(page);
30                 contentStream = new PDPageContentStream(doc, page);
31                 float scale = 0.5f;
32                 contentStream.drawXObject(ximage, 20, 400, ximage.getWidth()
33                         * scale, ximage.getHeight() * scale);
34                 
35                 PDFont font = PDType1Font.HELVETICA_BOLD;
36                 contentStream.beginText();
37                 contentStream.setFont(font, 12);
38                 contentStream.moveTextPositionByAmount(100, 700);
39                 contentStream.drawString("Hello");
40                 contentStream.endText();
41                 
42                 contentStream.close();
43             }
44         }
45         doc.save(file);
46         doc.close();
47     }

7、替換PDF檔案中的某個字串：

 1  /**
 2      * Locate a string in a PDF and replace it with a new string.
 3      *
 4      * @param inputFile The PDF to open.
 5      * @param outputFile The PDF to write to.
 6      * @param strToFind The string to find in the PDF document.
 7      * @param message The message to write in the file.
 8      *
 9      * @throws IOException If there is an error writing the data.
10      * @throws COSVisitorException If there is an error writing the PDF.
11      */
12     public void doIt( String inputFile, String outputFile, String strToFind, String message)
13         throws IOException, COSVisitorException
14     {
15         // the document
16         PDDocument doc = null;
17         try
18         {
19             doc = PDDocument.load( inputFile );
20 //            PDFTextStripper stripper=new PDFTextStripper("ISO-8859-1");
21             List pages = doc.getDocumentCatalog().getAllPages();
22             for( int i=0; i<pages.size(); i++ )
23             {
24                 PDPage page = (PDPage)pages.get( i );
25                 PDStream contents = page.getContents();
26                 PDFStreamParser parser = new PDFStreamParser(contents.getStream() );
27                 parser.parse();
28                 List tokens = parser.getTokens();
29                 for( int j=0; j<tokens.size(); j++ )
30                 {
31                     Object next = tokens.get( j );
32                     if( next instanceof PDFOperator )
33                     {
34                         PDFOperator op = (PDFOperator)next;
35                         //Tj and TJ are the two operators that display
36                         //strings in a PDF
37                         if( op.getOperation().equals( "Tj" ) )
38                         {
39                             //Tj takes one operator and that is the string
40                             //to display so lets update that operator
41                             COSString previous = (COSString)tokens.get( j-1 );
42                             String string = previous.getString();
43                             string = string.replaceFirst( strToFind, message );
44                             System.out.println(string);
45                             System.out.println(string.getBytes("GBK"));
46                             previous.reset();
47                             previous.append( string.getBytes("GBK") );
48                         }
49                         else if( op.getOperation().equals( "TJ" ) )
50                         {
51                             COSArray previous = (COSArray)tokens.get( j-1 );
52                             for( int k=0; k<previous.size(); k++ )
53                             {
54                                 Object arrElement = previous.getObject( k );
55                                 if( arrElement instanceof COSString )
56                                 {
57                                     COSString cosString = (COSString)arrElement;
58                                     String string = cosString.getString();
59                                     string = string.replaceFirst( strToFind, message );
60                                     cosString.reset();
61                                     cosString.append( string.getBytes("GBK") );
62                                 }
63                             }
64                         }
65                     }
66                 }
67                 //now that the tokens are updated we will replace the
68                 //page content stream.
69                 PDStream updatedStream = new PDStream(doc);
70                 OutputStream out = updatedStream.createOutputStream();
71                 ContentStreamWriter tokenWriter = new ContentStreamWriter(out);
72                 tokenWriter.writeTokens( tokens );
73                 page.setContents( updatedStream );
74             }
75             doc.save( outputFile );
76         }
77         finally
78         {
79             if( doc != null )
80             {
81                 doc.close();
82             }
83         }
84     }

上述描述的只是PDFBox的部分功能，在原始資源包中有很對例子，大家可以學習，PDFBox API路徑：http://pdfbox.apache.org/docs/1.8.8/javadocs/

轉載自：愛程式設計w2bc.com 如果本部落格無法幫助你，請看這裡，效果也很好：http://blog.csdn.net/w20228396/article/details/68065552

使用PDFBox處理PDF文件

【PDFBox下載】

PDFBOX處理PDF文件

使用PDFBox處理PDF文件

使用fileinput+pdfbox獲取pdf文件指定區域的內容

利用pdfbox將pdf文件轉換為圖片

C#操作PDF文件--PDFBox讀取pdf文件，O2S.Components.PDFRender4NET生成縮圖

PDFBox讀取PDF文件元資料

C#使用iTextSharp處理PDF文件

PDF文件解析：PDFBox和iText例項

Apache PDFbox開發指南之PDF文件讀取

Apache PdfBox 2.0.X 版本解析PDF文件（文字和圖片）

7.2 使用xpdf來處理中文PDF文件

[置頂] java處理office文件與pdf檔案(一)

Java使用PDFBox開發包實現對PDF文件內容編輯與儲存

LaTeX-WinEdt 編輯器和 PDF 文件的 Acrobat 11 程序關聯

刪除PDF文件中水印的方法

WPF中查看PDF文件 - 基於開源的MoonPdfPanel （無需安裝任何PDF閱讀器）問題匯總

python基礎—字符串處理、文件處理（運維必備）

從pdf 文件中抽取特定的頁面

怎樣打開與編輯PDF文件

分享將pdf文件轉換成圖片的圖文教程

使用PDFBox處理PDF文件

【PDFBox下載】

相關推薦