- Updates recommended RKorAPClient version to 1.4.1
- Requires Python >= 3.10 (support for Python 3.7–3.9 dropped, Python 3.14 added)
libv8is no longer needed as a system dependency- Python dicts are now converted to named R vectors (or named lists), e.g.
vc={"Germany": "pubPlaceKey=DE", "Switzerland": "pubPlaceKey=CH"}to label virtual corpora - Documentation and tests updated for the new features listed below, including labelling virtual corpora (
labelcolumn) and caching results withcacheAs - follows the DeReKo-KorAP-2026-II release:
corpusQuery()now fetches the newdmozDomainandwikiDomainmetadata fields, which replacedtextClass(RKorAPClient #31) - long running searches are no longer aborted after 10 minutes; only the connection's
timeoutapplies corpusQuery()andfrequencyQuery()now warn about failed or incomplete (timeExceeded) searches instead of silently reporting 0 or too few hits; their results carry a newqueryDurationcolumn- clearer verbose output: series of queries or virtual corpora are logged as an aligned table, with more accurate ETAs
cacheAsparameter for transparent result caching, now supported bycollocationAnalysis(),frequencyQuery(),corpusStats(),collocationScoreQuery()andtextMetadata()- names of a named
vcvector are consistently used as labels byfrequencyQuery(),corpusStats()andcollocationScoreQuery() - changed
logDicevalues: now computed as defined by Rychlý (2008), so values are comparable to Sketch Engine and other tools - changed
ll()values: contingency table now scales the sample size and both marginals with the window size (Evert 2004) collocationAnalysis()discards collocates occurring less often than expected by chance (newminObservedExpectedRatioparameter, default 1; set to 0 for the previous behaviour)- experimental support for comparing collocation analyses across multiple (optionally labelled) vcs
collocationScoreQuery()accepts a vector of collocates- fixed collocation analysis dropping snippets with certain markup shapes, e.g. sentence-bounded
contains(<base/s=s>, ...)queries (RKorAPClient #14) - fixed score threshold in recursive collocation analysis
- Updates recommended RKorAPClient version to 1.2.1
- fetchAnnotations() method added to KorAPQuery class, to fetch annotations for all collected matches
- Updates recommended RKorAPClient version to 1.1.0
- fixed bug with fetching result pages with an offset >= 10,000 (=1e+05 ...) issue #25
- timed out corpus queries are no longer cached (see issue #7)
- improved stability of
cifunction - improved error handling
- improved logging
- added ETAs to logging in verboose mode
- Fixed NULL warning in fetch functions
- Simplified authorization process for accessing restricted data via the new
auth()function - Fixed issues with tokenized matches in
corpusQueryresults - Fixed smoothing constant in
mergeDuplicateCollocatesfunction - Fixed chainability of fetch methods in
corpusQuery
- Updates recommended RKorAPClient version to 0.9.0
- Added
matchStartandmatchEndcolumns to corpusQuery results, containing the start and end positions of the match in the text - Added
mergeDuplicateCollocatesfunction to merge collocation analysis results for different context positions - Added a query column to collocation analysis results
- Improved documentation for span parameter in
collocationAnalysisfunctions - Updated
textMetadatamethod to use new metadata fields API, if available, to retrieve custom metadata for a text based on its sigle - Added new unit tests to cover the new features and changes
- Updates recommended RKorAPClient version to 0.8.1
- fixed rare frequencyQuery incompatibility with outdated KorAP instances (see KorAP/Kustvakt#668)
-
Updates recommended RKorAPClient version to 0.8.0
-
Added
textMetadataKorAPConnection method to retrieve all metadata for a text based on its sigle -
Added
webUiRequestUrlcolumn also to corpusStats results, so that also virtual corpus definitions can be linked to / tested directly in the KorAP UI -
Uses server side tokenized matches in collocation analysis, if supported by KorAP server
-
Unless
metadataOnlyis set, also tokenized snippets are now retrieved in corpus queries (stored inres.slots['collectedMatches']['tokens.left'],res.slots['collectedMatches']['tokens.match'],res.slots['collectedMatches']['tokens.right']). Because Pandas data frames cannot store lists, tokens are stored as strings, tab separated. -
Python 3.11 and 3.12 are now supported
-
Python 3.7 support has been dropped (by rpy2 dependency)
- Updates recommended RKorAPClient version to 0.7.5
- fixes collocation scores for lemmatized node or collocate queries
- Automatically converts again between rpy and py objects, most importantly between Pandas and R data frames, also with newer versions of rpy2
- Fixes "Hello world" example in Readme.md
- Updates references / citation information
- Changes corpusStats to return a pandas.DataFrame by default
- Advertises Python 3.10 as supported
- Adds interactive plot examples using Vega-Altair