jericho解析html
阿新 • • 發佈:2017-07-08
jericho解析html
1.導入jar包
2.實現源代碼
package com.zhishang.lucene; import net.htmlparser.jericho.Element; import net.htmlparser.jericho.HTMLElementName; import net.htmlparser.jericho.Source; import org.junit.Test; import java.io.File; import java.io.IOException; /** * Created by Administrator on 2017/7/8. */ public class HtmlBeanUtil { @Test public void parseHtml(){ String path = "G:\\data\\index.html"; try { Source sc = new Source(new File(path)); Element element = sc.getFirstElement(HTMLElementName.TITLE); System.out.println(element.getTextExtractor().toString()); System.out.println(sc.getTextExtractor().toString()); } catch (IOException e) { e.printStackTrace(); } } }
本文出自 “素顏” 博客,請務必保留此出處http://suyanzhu.blog.51cto.com/8050189/1945451
jericho解析html