Quick fix google scholar entry fetching #2082

matthiasgeiger · 2016-09-28T10:05:20Z

Google Scholar fetching was broken again (see #1886)

With this fix at least getting the first 10 search hits is possible again.

Configuration is no longer possible in the current form and google generally limits the responses (per page) to 20 hits (however, even using this will cause a captcha challenge for JabRef).

As only 10 hits are allowed a rewrite to the new FetcherInfrastructure should now be possible (thus, omitting the 2-step approach).

Change in CHANGELOG.md described
Tests created for changes
Screenshots added (for bigger UI changes)
Manually tested changed features in running JabRef
Check documentation status (Issue created for outdated help page at help.jabref.org?)
- Update of http://help.jabref.org/en/GoogleScholar might be useful to indicate that only 10 hits are (currently) shown

tobiasdiez

Some small remarks.

tobiasdiez · 2016-09-28T13:25:40Z

src/main/java/net/sf/jabref/gui/importer/fetcher/GoogleScholarFetcher.java

        int lastRegionStart = 0;

        while (m.find()) {
-            String link = m.group(1).replace("&amp;", "&");
-            link = link+"&oe=utf-8"; // append param 'oe=utf-8' to tell google to serve UTF-8 encoded results


The UTF-8 string is no longer required?

Yes, as I am using a Firefox Useragent UTF-8 is automatically delivered.

tobiasdiez · 2016-09-28T13:27:29Z

src/main/java/net/sf/jabref/gui/importer/fetcher/GoogleScholarFetcher.java

-            String link = m.group(1).replace("&amp;", "&");
-            link = link+"&oe=utf-8"; // append param 'oe=utf-8' to tell google to serve UTF-8 encoded results
+
+            String citationsPageURL = CITATIONS_PAGE_URL_BASE+m.group(1)+CITATIONS_PAGE_URL_SUFFIX;


Please use Apache's URIBuilder instead of String concatenation.

Mhmm. I don't think this is that much an improvement as still some string concatenation is needed for the param "q" with "info:"+id+":scholar.google.com/"

tobiasdiez · 2016-09-28T13:34:02Z

src/main/java/net/sf/jabref/gui/importer/fetcher/GoogleScholarFetcher.java

+
+            String citationsPage = URLDownload.createURLDownloadWithBrowserUserAgent(citationsPageURL).downloadToString(StandardCharsets.UTF_8);
+
+            Matcher citationPageMatcher = GoogleScholarFetcher.BIBTEX_LINK_PATTERN.matcher(citationsPage);


It is not possible to get the BibTeX Url directly from the id?
For example, directly access https://scholar.googleusercontent.com/scholar.bib?q=info:b2pGeL14LLMJ:scholar.google.com/&output=citation&scisig=AAGBfm0AAAAAV-vITSVVNk7OSER8S_LFaMSElM9jnUcv&scisf=4&ct=citation&cd=-1&hl=en

No, as this "scisig" is necessary and seems to be changing.

That's unfortunate.
I did some research and I think I found something.
Google stores the setting to display bib-links in a cookie. Sending the cookie GSP=IN=d192d757fd09588a+7e6cc990821af64:CF=4 along the first query directly shows the Import into Bibtex link,

<a href="https://scholar.googleusercontent.com/scholar.bib?q=info:b2pGeL14LLMJ:scholar.google.com/&output=citation&scisig=AAGBfm0AAAAAV-vafm6NWi20weGxxou9W2xi8GzZ8YCf&scisf=4&ct=citation&cd=0&hl=en" class="gs_nta gs_nph">Import into BibTeX</a>

Appearently, the part after GSP=IN=... is not that important since my real id had 63 at the end instead of 64.

Ah using a Cookie instead of the config method can do the trick...

I'll merge the current state in and than start implementing a new version based on the new fetcher interfaces. Using these we should be able to perform automated tests which will be helpful for the initial development and to detect issues caused by changes of Google structure faster...

fix google scholar entry fetching

43134d9

matthiasgeiger added this to the v3.7 milestone Sep 28, 2016

matthiasgeiger added status: ready-for-review Pull Requests that are ready to be reviewed by the maintainers component: search labels Sep 28, 2016

matthiasgeiger mentioned this pull request Sep 28, 2016

Google search issue in version 3.6 [Fixed in DevBuilds] #1886

Closed

stefan-kolb approved these changes Sep 28, 2016

View reviewed changes

tobiasdiez requested changes Sep 28, 2016

View reviewed changes

matthiasgeiger merged commit f905e29 into master Sep 28, 2016

matthiasgeiger deleted the fix-googlescholar branch September 28, 2016 17:59

matthiasgeiger mentioned this pull request Sep 30, 2016

Rewrite google scholar fetcher to new infrastructure #2101

Merged

5 tasks

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Quick fix google scholar entry fetching #2082

Quick fix google scholar entry fetching #2082

matthiasgeiger commented Sep 28, 2016

tobiasdiez left a comment

tobiasdiez Sep 28, 2016

matthiasgeiger Sep 28, 2016

tobiasdiez Sep 28, 2016

matthiasgeiger Sep 28, 2016

tobiasdiez Sep 28, 2016

matthiasgeiger Sep 28, 2016

tobiasdiez Sep 28, 2016

matthiasgeiger Sep 28, 2016


		String citationsPage = URLDownload.createURLDownloadWithBrowserUserAgent(citationsPageURL).downloadToString(StandardCharsets.UTF_8);

		Matcher citationPageMatcher = GoogleScholarFetcher.BIBTEX_LINK_PATTERN.matcher(citationsPage);

Quick fix google scholar entry fetching #2082

Quick fix google scholar entry fetching #2082

Conversation

matthiasgeiger commented Sep 28, 2016

tobiasdiez left a comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment

Choose a reason for hiding this comment