Every scientific publication and research material indexed against SHARE (Survey of Health, Ageing and Retirement in Europe). Semantic retrieval finds work on health, ageing, retirement, cognition and socio-economic conditions across European populations even when it uses different vocabulary from your query.
SHARE publication index
Upload a SHARE paper
Submit a PDF that you (co-)authored or have rights to share. Metadata is filled in from the DOI and checked against the existing index. The PDF is held inside SHARE only — it is not redistributed.
About this service
The SHARE Research Portal is a publication-discovery and research-support environment for the SHARE (Survey of Health, Ageing and Retirement in Europe) community. It brings together semantic search, a browsable publication index, topic-based exploration, and SARA — the AI research assistant that works over the SHARE-related scientific literature.
Who runs it, and what it answers from
SARA is built and run by SHARE Austria and runs on SHARE's own servers. It answers from the publication index, the released microdata and the codebooks — not from the open internet.
The model is Qwen3.8-Flash-Next, running on this machine, so no outside AI service reads what you type or what it reads on your behalf. A search for papers outside SHARE's collection sends its search words to the public publication indexes, which the tool list beside the message box names. It holds up to 524,288 tokens of context, which is several long reports and the conversation around them, so a long working session does not quietly lose its earlier half. The same model does the rewrite behind the wand beside Send.
One assistant, and what your account can reach
There is one SARA. What differs is the set of tools your account carries. A researcher account can search the literature, look variables up, ask for aggregate figures and read files you attach. A SHARE staff account carries those and two more, the web and SHARE's own copy of a paper, because staff use the same assistant for correspondence and admin as for research.
There is no mode to switch and no setting to find. The tool list beside the message box shows exactly what is available in the conversation you are in.
Help us improve
Conversations are recorded anonymously and used to improve the assistant. Rating an answer with the thumbs is useful; saying what was wrong is far more useful. One line of correction is worth more to us than a hundred silent downvotes.
Getting help
Anything that looks broken rather than merely wrong belongs in the support inbox. The service status pill in the top bar shows whether the problem is yours or ours.
What it does
- Research assistant (SARA): conversational access to the publication corpus with citation support, summarisation, and methodology questions grounded in the SHARE literature — and to the SHARE microdata, for questions about variables, cross-wave harmonisation and analytical workflows. There is one SARA and no mode to choose: she has the literature tools and the data tools in the same conversation, and which further tools she can reach follows from your account rather than from anything you switch.
- Rewrite a draft (the wand beside Send): type or paste something you have already written, press the wand, and pick a style — journal article, project report, policy brief, internal note, or describe the register you want in your own words. The rewrite is done by the same model that answers you, and it comes back in the conversation so you can read it before you do anything with it. It edits wording, not facts: it does not check whether a claim is true, look anything up, or add to what you wrote — so read it against your original. Nothing leaves this service, and nothing is sent anywhere until you send it.
- Search publications: semantic and hybrid (BM25 + dense vector) search over the full text of indexed publications, with optional neural reranking.
- Browse by topic: OpenAlex-taxonomy-based topic tree (domain → field → subfield → topic) with publication timelines.
- Publication index: bibliographic inventory compiled from the SHARE repository crawler, including entries whose full text is available and those known by reference only.
Semantic search — what it is, and why to use it
Traditional keyword search only finds papers that contain the exact words you type. Semantic search goes a step further: queries and paper passages are both converted into high-dimensional numerical vectors (embeddings) that capture the meaning of the text rather than its surface form. Two passages that talk about the same concept end up close together in that vector space even if they share no vocabulary.
This has several concrete advantages for literature discovery:
- Paraphrase and synonym tolerance: a query for “cognitive decline in old age” also surfaces papers written as “neurocognitive ageing trajectories” or “memory loss in the elderly”.
- Natural-language queries: you can ask full questions (“how does early retirement affect depression risk?”) instead of guessing the right keyword boolean.
- Multilingual robustness: because the embedding model places semantically related terms close together, near-synonyms and closely related concepts across languages or disciplines tend to match.
- Better recall on conceptual questions: semantic search finds relevant material that keyword search would miss because the authors used different terminology.
The portal uses a hybrid retrieval strategy: semantic (dense-vector) scoring is combined with classical BM25 keyword scoring, so you keep the precision of exact term matches (for example specific variable names, author names, or waves) while also benefiting from meaning-based recall. An optional neural cross-encoder reranker then re-scores the top candidates by looking at the query and each passage together, which sharpens the ordering of the final results.
Data sources
- SHARE document repository (scientific publications, working papers, methodology reports, questionnaires, codebooks).
- Bibliographic inventory produced by the SHARE repository crawler
(
repository_inventory.bib), enriched with DOI and OpenAlex-topic assignments where available. - Topic assignments derived from the OpenAlex topic taxonomy; papers without an OpenAlex match receive zero-shot assignments as a fallback.
How it works
- Text extraction: PDF content is extracted, cleaned, and chunked before indexing.
- Embeddings: chunks are embedded with
BAAI/bge-m3— a multilingual model, chosen because the SHARE corpus is not written only in English — and stored in a FAISS vector index. - Hybrid retrieval: queries are answered by combining
BM25 lexical scoring with dense-vector similarity; an optional
cross-encoder reranker (
BAAI/bge-reranker-large) refines the top results. - Research assistant: SARA runs on Qwen3.8-Flash-Next, hosted on this server, with a 524,288-token context window. It answers with retrieval against the publication index, the released microdata and the codebooks, and cites the sources it draws from.
- Checking a draft questionnaire item: SHARE's 5,864 source-language questions are embedded with the same model, so an item being written can be matched against everything already fielded — which waves asked it, how the wording moved, and what changing it would cost in comparability.
- Topic exploration: per-paper topic assignments are aggregated into a domain/field/subfield/topic tree and visualised as a publication-per-year timeline.
Index refresh
The index is refreshed on a regular schedule. New publications added to the SHARE repository are picked up by the next indexing pass, which reconciles the docstore against the live repository, re-runs text extraction and embedding for new files, and rebuilds the topic assignments.
Access
Access to the research portal is limited to registered SHARE-ERIC data users in good standing. Use of the service is governed by the SHARE-ERIC Conditions of Use and is restricted to scientific research consistent with each user's registered SHARE project.
Technical stack
- Vector store
- FAISS
- Embeddings
BAAI/bge-m3— multilingual, 1024 dimensions- Lexical index
- BM25
- Reranker
BAAI/bge-reranker-largecross-encoder- Back end
- FastAPI (Python)
- Front end
- Static HTML, CSS and JavaScript — no build step
- Assistant model
- Qwen3.8-Flash-Next, 524,288-token context
- Rewrite model
- Qwen3.8-Flash-Next — the same model, for the wand beside Send
- Assistant inference
- Locally hosted, on this server. Nothing is sent to a third party
- Topic taxonomy
- OpenAlex, with zero-shot fallback
Using SARA for data work
There is nothing to switch on and no mode to choose. Ask about the microdata in the same conversation you ask about the literature, and change subject mid-conversation if that is where the work goes — which it usually is, since a question about a variable and a question about the paper that used it are the same question from two sides.
What that looks like in practice:
- Finding the variable you mean from a description of it, rather than from its name — and reading the question as it was actually asked, in any of the 42 fielded languages.
- Checking comparability before you pool waves. Wording moves between waves, and SARA shows you where it moved rather than asserting that a series is safe to combine.
- Running the analysis. Python and R execute on this server against the release, and the code that produced a number is shown alongside it, so you can check the working rather than trust the answer.
- Working on your own files. Attach a spreadsheet or a document and she will read it, work on it, and hand back a file.
One limit is worth knowing before you start. Results are aggregate — counts, means, distributions, models — and never a single person's record. What you upload stays on this server. The files of a conversation, yours and the ones she makes, are in the Files panel, and you download them from there. Out of the secure environment a file leaves only through an export request.
No setup beyond being signed in. Which further tools SARA can reach follows from your account, and the tool list beside the message box states it exactly.
Connecting external AI tools (MCP)
SHARE is exposed as a Model Context Protocol server, so external assistants (Claude.ai, ChatGPT, Claude Desktop, any MCP client) can query the literature and the survey documentation mid-conversation — and get back records rather than a plausible-sounding guess.
What it answers:
search_share_variables- Which variable measures a construct — across 57,679 variable labels and the wording of the questions behind them.
get_share_question- What a question actually asked, as fielded in any of 42
language versions — including regional variants such as
CH-de,BE-fr,IL-arandLU-pt. Your model never has to translate SHARE wording, because it can ask for the original. compare_share_waves- Whether an item was worded the same way across waves — the check that should happen before anyone pools them.
search_papers,get_topic_tree,get_topic_timeline,search_by_topic- The publication index and its OpenAlex topic taxonomy.
Metadata only. No microdata is reachable through this server — not filtered, not aggregated, not one row. That is a property of what the endpoint can execute, not a policy it promises to follow. Analysis stays inside the portal, where you are signed in and the data never leaves.
Two habits are built in. Every record states which documentation release it came from, because documentation and data are not always the same vintage. And where something is not documented, the answer says so and returns what is — constructed variables and national add-on questionnaires sit outside the central instrument, and a label is not a question.
Connection details:
- MCP server URL
https://research.share-austria.at/mcp- Authentication
- OAuth 2.0 (authorization code + PKCE)
- OAuth client ID
- your own portal Client ID — the same one you sign in with above.
- OAuth client secret
- your own portal client secret. If you have mislaid it, ask the portal administrator to reissue one; it is stored only as a hash and cannot be looked up.
Use your personal credentials, never a shared one. Every connection is attributable to one account, which is what makes access revocable for a single user without disrupting anyone else, and what keeps the usage metrics meaningful. Do not paste your secret into a shared machine, a group chat, or a config file you commit to a repository.
Claude.ai (web):
- Open claude.ai/settings/connectors and choose “Add custom connector”.
- Paste the server URL, then your own portal Client ID and client secret.
- Choose Connect; you are redirected through the OAuth flow and back to Claude.
- Ask, for example, “which SHARE variable measures life satisfaction, what exactly was asked in Italian, and was it worded the same way in wave 4?” — three tool calls, three records, no paraphrase.
ChatGPT: under Settings → Integrations (or Custom Actions), choose “Add MCP server”, paste the same credentials, and complete the OAuth authorisation.
Claude Desktop: add the following to your
claude_desktop_config.json
(~/Library/Application Support/Claude/ on macOS,
~/.config/Claude/ on Linux,
%APPDATA%\Claude\ on Windows), then restart the app:
{
"mcpServers": {
"share-research-portal": {
"url": "https://research.share-austria.at/mcp",
"auth": {
"type": "oauth2",
"client_id": "YOUR-PORTAL-CLIENT-ID",
"client_secret": "YOUR-PORTAL-CLIENT-SECRET"
}
}
}
}
External MCP clients are subject to the same SHARE-ERIC Conditions of Use as the portal itself — only registered SHARE users in good standing may connect, and all queries are logged with the same privacy-safe usage metrics as direct portal traffic.
Using SARA's model in your own tools (OpenAI-compatible API)
The model behind SARA also answers at an OpenAI-compatible address, so a coding assistant in your editor (Zoo Code, Continue), a script that uses an OpenAI library, or any other client that speaks that protocol can use it with your portal login. The model runs on this server; your requests do not go to an outside AI provider.
Connection details:
- Base URL
https://research.share-austria.at/openai- API key
- your own portal client secret, the one you sign in
with above (it begins
psc_). The authenticator code is not needed here. - Model
sara- Context window
- 524,288 tokens
When your client brings its own tools, as an editor does to read and change
your files, the model answers and your client carries out the steps on your
computer. Nothing on your computer is reachable from this server. A request
that sets "tools": true instead gets SARA's own tools, the ones
your account has in the portal, working in your own files here.
The model is shared with everyone who uses SARA, so when it is busy a request waits for its turn.
Zoo Code (VS Code): in its settings choose the API provider
OpenAI Compatible, and enter the base URL, the API key and the model
above. Set the context window in the model's settings to
524288, or Zoo Code assumes a smaller one.
Continue (VS Code, JetBrains): add the model to your
config.yaml:
models:
- name: SARA
provider: openai
model: sara
apiBase: https://research.share-austria.at/openai
apiKey: YOUR-PORTAL-CLIENT-SECRET
roles:
- chat
- edit
- apply
Python, with the openai package:
from openai import OpenAI
client = OpenAI(
base_url="https://research.share-austria.at/openai",
api_key="YOUR-PORTAL-CLIENT-SECRET",
)
reply = client.chat.completions.create(
model="sara",
messages=[{"role": "user",
"content": "What does a fixed-effects model control for?"}],
)
print(reply.choices[0].message.content)
The rules for MCP above apply here too. Use your own credentials, never a shared one, and keep them out of any file you commit to a repository; an editor's settings file is easily one of those. Use it for your work with SHARE, under the same SHARE-ERIC Conditions of Use as the portal. Every request is attributed to your account.
Contact
Questions, bug reports, or feedback can be directed to the SHARE Austria research support team.
Legal note — applicability of the EU AI Act
SARA falls outside the scope of Regulation (EU) 2024/1689 (the AI Act) under the research exemption in Article 2(6): “This Regulation does not apply to AI systems or AI models, including their output, specifically developed and put into service for the sole purpose of scientific research and development.”
SARA was purpose-built for SHARE's research infrastructure. It is not an adaptation of a general-purpose product, it is not offered commercially, and it is not used outside the SHARE/JKU research context. It indexes SHARE's own research data and supports the research process directly — variable search, data work, and drafting support tied to SHARE outputs.
Recital 25 keeps a tool that is merely used for research inside the scope of the Regulation. That carve-back is aimed at incidental use of general-purpose products — a researcher reaching for a commercial chatbot — not at infrastructure purpose-built for a specific research programme. SARA sits on the “specifically developed” side of that line, not the “happens to be used by researchers” side. The accepted example of this exemption — an AI instrument used within a scientific research programme, such as a neuroscience measurement tool — is structurally the same case: SARA is an AI instrument within SHARE's research programme, not a research project about AI, and the exemption is not limited to research on AI itself.
The position depends on scope, and the scope is held deliberately. SARA's use stays within SHARE's research activity; drift into general administrative or IT-support use unrelated to the research would weaken the “sole purpose” claim, and is either avoided or handled through a separate, clearly delineated channel.
Access and use conditions
SARA is a research aid, not an author. Responsibility for what you publish stays with you, as does responsibility for the disclosure rules that apply to SHARE data. An answer that cites a paper is not the same as having read it.
What stays on the server
Microdata, free text and anything you upload stay on SHARE's servers. Aggregate figures may be taken out once they pass a disclosure check — minimum cell counts, dominance rules, no listing of individuals. An export is a request, not a download button.
Staff accounts additionally have a web tool. It is the only thing here that reaches off SHARE's network, and no conversation content or SHARE data is sent to it.
What is recorded
Conversations are recorded anonymously and used to improve the assistant. Deleting a conversation deletes it for us as well; there is no separate training copy. Retention and your rights over that record are set out in the privacy notice.
Entitlement
Using the portal and SARA costs nothing. They are for registered SHARE users and SHARE staff; nobody else has access.
Use of SARA follows the portal's conditions of use, and access to the released microdata follows your SHARE Research Data Centre registration. Neither is granted by this page.
Conditions of use
Access to the research portal and associated AI services is provided exclusively to registered SHARE-ERIC data users in good standing.
Use of this service and any AI outputs is subject to and governed by the SHARE-ERIC Conditions of Use and is strictly limited to scientific research purposes consistent with each user's registered SHARE project and user licence.
No additional rights are granted beyond those already conferred under the SHARE data access framework.
Note on scientific research use
Under EU copyright law, research organisations may carry out text and data mining (TDM) of works to which they have lawful access, for the purposes of scientific research (Article 3 of Directive (EU) 2019/790; in Austria, § 42h (1) UrhG). Rightsholders cannot opt out of this exception, and contract terms that try to exclude it are void (Article 7(1)). SHARE-ERIC, as a non-profit research infrastructure, relies on it.
Lawful access clearly includes subscriptions, licences and open-access publications. Whether it extends to everything that can be found freely online is not yet settled.
German courts have upheld the research exception at two instances in Kneschke v. LAION: the Hamburg Regional Court in 2024 and the Hanseatic Higher Regional Court in 2025. A further appeal is pending before the Federal Court of Justice, and a referral to the Court of Justice of the EU is likely. This note reflects the position as of 25 September 2026.
The exception covers copying and analysing works for research. It does not cover making them available to others: reproducing source texts verbatim in what you publish or share remains subject to copyright as usual.
The research portal and associated AI services form part of SHARE-ERIC's controlled research infrastructure. Access is therefore limited to registered SHARE users and may only be used for scientific research in accordance with the SHARE Conditions of Use.
Use of the research portal or AI services for commercial purposes, general public services, or other non-research activities is not permitted.
Usage metrics and data collection
To operate and improve service reliability, SHARE-ERIC collects privacy-safe, aggregated usage metrics for the research portal and AI endpoints.
Collected metrics include:
- Service and endpoint name
- Time bucket (hour or day) and total request counts
- HTTP status classes (2xx/4xx/5xx) and latency percentiles
- For search endpoints only: query-length buckets and result-count buckets
Not collected in these metrics:
- Raw query text, prompts, or model responses
- Uploaded document contents
- API keys or authentication secrets
- Full IP addresses or direct personal identifiers
Public metrics views apply low-volume suppression (a k-anonymity threshold) to reduce re-identification risk.
Secure Research Environment
Somewhere to work with SHARE data where the data does not move. R and Python, a working directory of your own, and SARA alongside — all of it on SHARE's servers, in a browser tab.
What it is for
Analysis of SHARE data that you would otherwise have to download first. You open the environment in a browser, write and run your code there, and the data stays where it is. Nothing is copied to your own machine, and nothing needs to be.
Your working directory is yours alone, and it is encrypted. It is bounded by space rather than by time, so a project that runs for four years is not interrupted by a clock.
How results leave
Copy a file into the export/ directory of your working directory. That copy is the request: it is picked up within seconds, checked, and — if it may leave — appears in the portal under My exports for you to download. The answer is also written back to export/receipts/, so you can stay where you are working.
Two things do not leave: the survey's records, and anything that points at one person. Everything else is research output and is meant to leave. In practice that means tables, model output, figures and documents pass; record-level extracts do not, and neither do statistical data files (.dta, .sav, .RData, .rds, .parquet) or archives in any form — those are how a dataset travels rather than how results do.
A refusal says what would need to change. Fix the file, drop it again, and you have another answer in seconds; if you think the answer is wrong, ask a person from the My exports page and it becomes a support ticket.
Tables, model output, plots, summary statistics.
Record-level extracts, free-text answers, anything identifying a respondent.
R, Python and SARA
R 4.6.1 and Python 3.10, with a pinned package snapshot — an analysis run today and rerun in three years resolves the same versions. It is a complete development environment: debugger, version control, terminal, package management. The restriction is on what leaves, not on what you can do.
SARA runs on the same machine as the data rather than in the cloud. It can read your script, explain what a variable means, and propose a change for you to accept or reject.
On Stata: SHARE data in Stata format is fully usable — haven reads .dta files including value labels. What cannot run here is an existing .do file, because Stata itself is not installed. If your work is in Stata, the data comes across; the code has to be rewritten.
Getting access
Being signed in to the portal is not enough. SHARE Austria grants access on application, and checks it in advance rather than at the point of use. You will be asked to confirm you are registered with the SHARE Research Data Centre, where you work, and what you intend to do with the data.
That last part is what the decision turns on. A few specific sentences — the question, the waves and countries, roughly what analysis — are worth more than a page of generalities.
Apply for access
One person at SHARE Austria reads this. The first part is verification; the second is the part the decision turns on.
Endpoint usage
Aggregated, privacy-safe request metrics for the research portal and its AI endpoints. No query text, prompts, model responses or direct identifiers are recorded.
Endpoint overview
| Service | Endpoint | Requests | 2xx | 4xx | 5xx | P50 (ms) | P95 (ms) | P99 (ms) | Notes |
|---|
Timeseries
| Bucket | Service | Endpoint | Requests | 5xx | 5xx rate | Notes |
|---|
Buckets below the k-anonymity threshold are suppressed, so low-volume endpoints are omitted rather than reported.