Examining MetaGer's Merged Results From Behind a Proxy
MetaGer is a meta search engine: it passes your query to multiple sources, merges and deduplicates the returned lists, and shows you a single page. This structure makes regional analysis through a proxy different from single-index engines. This page explains where that difference comes from and how to measure it.
Aggregation logicHow does reducing multiple sources into a single list change the outcome?
02
Freshness chainIs the delay in the source's crawl schedule or in the merging layer?
03
Operator forwardingWhere does an operator get lost as it is passed on to the sources?
04
Regional exitWhich layer the exit country touches, and which it doesn't.
In meta search engines the results page is not a snapshot of a single index. Your query is sent in parallel to multiple sources, the returned lists are merged, records pointing to the same address are deduplicated, and the remaining set is arranged into a single ranking. The ordering you see is the output of that merging decision — it exists in exactly that form in none of the sources.
This has a concrete consequence for proxy work. When you change your exit country, the effect originates not in the merging layer but in the sources' own responses. If a source is sensitive to regional signals its contribution changes; an insensitive source's contribution stays the same, and in the merged list you see the sum of the two. That is why the difference is usually softer and less predictable than with a single-index engine.
The sections below unpack this chain in order: the merging logic, freshness, operator forwarding, the exit-type decision, the verification method, the regional view and finally setup and limits.
How is the merged results page built?
The chain starts with your client. The query goes to MetaGer's own endpoint; from there it is distributed simultaneously to several sources. Each source returns a list with its own ranking. The aggregator takes these lists, collapses records pointing to the same target into a single row, and produces the final order with its own weighting. Behind every row you see on the results page there is at least one source, and behind some of them more than one.
The proxy is involved only in the first step of this chain: the connection between you and the aggregator. How the aggregator connects to the sources is independent of your network setting. This distinction also defines a limit — your exit country takes effect on the sources' side not as a direct geographic signal, but through the context the aggregator passes along. So when you change country, don't expect as sharp a picture as you would see on a single-index engine.
The second practical consequence concerns diagnosis. If a result isn't where you expected, there can be two different reasons: none of the sources returned that page, or one did but it was pushed down during merging. To tell these apart, repeat the same query on a single-index engine; an engine such as Mojeek, which relies on its own crawl, provides a good reference point, and you can reach its guide from the related pages list at the bottom of this page.
Note
If the results page indicates which source a record came from, write that information into your comparison notes too. Source distribution is the strongest clue as to where the difference between two runs came from.
DIAGRAMThe nodes stretching from the query to the merged list
You can scroll the diagram horizontally to inspect it
The proxy is involved only on the edge between the client and the aggregator; the connections to the sources lie outside your network setting.
The freshness chain: where does the delay accumulate?
On a merged page, currency is not tied to a single schedule. Each source has its own crawl frequency; a page may already be in one source's index while it hasn't yet entered another's. In the merged list this shows up as the page appearing, but sitting lower than you expected. The slowest link in the chain sets the practical freshness limit for that query.
The proxy adds nothing to this picture. Changing the network path affects neither the sources' crawl schedules nor the aggregator's own refresh interval. The observation that "results look newer when I view them from another country" is almost always the consequence of something else: a different weighting of the source mix on that request, or a cache on your side.
To account for caching on your own side, it is enough to open each run in a clean window with caching disabled; the mechanism is described in detail in the HTTP proxy caching article. The point that really demands attention is specific to merged pages and is very easily confused with caching: when a source is slow to respond, the aggregator does not wait for it indefinitely — it builds the page from the lists it has. The same query, the same exit and the same minute can produce two different source mixes.
This behaviour looks like a freshness difference, but it is a timeout behaviour and it can be distinguished. Repeat a slow run exactly as it was. If the missing source returns to the list on the second run, what you saw was not index currency but a source that dropped out on the first run; if it doesn't return, that source genuinely isn't contributing for that query. Claiming "results are newer when viewed from that country" without making this distinction is the most easily refuted claim in the whole measurement.
When looking for a new page, wait at least a few hours; its absence is not a regional difference.
Open every run in a clean profile; disable the cache.
Note how many sources contributed on each run; a missing source looks like a freshness difference.
Don't build a freshness claim without repeating the slow run.
How are operators passed from the aggregator to the sources?
On a single-index engine an operator is interpreted in one place. In meta search there are two stages: the aggregator parses the operator and forwards it to the sources; each source then interprets it by its own rules. If a source supports exact-phrase search its contribution narrows, while a source that doesn't continues to return a broad list. On the merged page you see the sum of the two — which is why, when you use an operator, the narrowing may not be as sharp as you expect.
Testing this behaviour isn't hard. Run the same query in three forms: plain, as an exact phrase, and with one word excluded. Note the result count and the top ten records on each run. In the exclusion test, make sure the word you exclude actually appears in the results of the plain run; otherwise what you are measuring is not the operator but an empty change.
Focus or category options behave like operators too: they route the query to a different set of sources. When you move from web results to a news- or academic-focused set, your comparison resets, because you are no longer looking at the same sources. Keep the focus fixed throughout the measurement and write down which focus you were working in.
Stage
What it does
What it means for measurement
Query parsing
Recognises the operator and prepares it for forwarding
A typo quietly turns into plain text here
Distribution to sources
Forwarded to each source in its own format
Support varies from source to source
Response merging
Lists are deduplicated and ranked
The narrowing is smoothed out in the total
Focus selection
Changes the set of sources
The comparison resets — keep it fixed
Thinking about the exit-type decision along two axes
In work that involves reading search pages, two axes determine the exit choice. The first is pool diversity: across how many different networks and regions your addresses are spread. The second is session stability: how long a single address stays with you. These two properties often move in opposite directions, and the right choice depends on the kind of work.
In comparative measurement, stability generally matters more than diversity, because knowing that every run came from the same address is what makes the picture defensible. For that reason, in work spread over several days an ISP proxy or a static datacenter exit is preferred. If, on the other hand, you want to collect a broad sample from different regions, diversity takes priority and a residential pool is the better fit.
Distribution matters as much as pool size: addresses spread across many different subnets are more resilient than a large pool packed into a single block; when one subnet is flagged, the entire pool doesn't lose value at once. In work that calls for rotation, a pool that changes the address on every request spreads the load; just don't neglect to record which run came from which address, or the comparison becomes untraceable.
DIAGRAMWhere exit types sit on diversity and stability
You can scroll the diagram horizontally to inspect it
The positions are a relative arrangement, not a measurement: the top-right corner is not ideal, merely the region where both properties are high together.
Your exit plan for MetaGer analyses
For comparisons spread over several days a static address is preferred; for collecting a broad sample, a pool with high diversity.
Choose whichever you need from our residential proxies, datacenter proxies, IPv6 and ISP solutions. Every plan comes with unlimited options, 99.9% uptime, rotating proxies, sticky sessions and 24/7 support. Ideal for web scraping, ad verification, SEO monitoring and digital data collection.
ISP ProxyStatic Turkish IPs registered to an ISP
ISP-registered static Türkiye IPs; they combine datacenter speed with the reputation of a real carrier. Ideal for long sessions and low-ping use.
On a meta engine, saying "results are different in that country" requires more careful evidence than on a single-index engine. Because in a merged list even a small shift in weighting can move the ranking, and that looks like a regional difference. The general sampling method — how the query set is fixed, how runs are spread across the calendar, why a control group is needed — is explained in the sampling section on the Swisscows page . Here we look only at the evidence problem specific to merged pages.
That problem is this: in meta search a record's identity has two parts. Its address on one side, and which source put it into the list on the other. If the results page shows the source label, store that label together with the addresses of the top ten rows; if it doesn't, at least note how many different sources contributed. That way you end up not with a single ranking list but with two series that check each other.
Run the comparison across these two series, in this order. First look at the source distribution: has the number of rows each source contributed changed between the two runs? If it has, you cannot read a shift in ranking as a regional difference, because the raw material entering the list has changed. If the distribution is constant but the ranking moves, the difference lies in the merging weights and is the engine's own decision, not your exit's. Only in the third case — distribution constant, weighting the same, and yet different addresses coming back — do you have a defensible regional finding.
If you are going to automate the measurement, keep the request pace close to human usage; if you are hitting a rate limit, the solution is not to increase the number of addresses but to widen the interval between runs and shrink the query set. In meta search this matters even more, because a single query creates work for several sources in the background and the page waits for the slowest one. When you compress the pace, what you measure is not the engine's behaviour but the queue you built up yourself. The practical principle is simple: keep the number of concurrent connections below the limit your provider defines, and put a fixed wait between runs.
Tip
At the start of each run, read the exit address with my IP address and write it on the first line of your log. When the sticky window expires mid-run, that line is the only evidence that tells you whether a shift in source distribution came from the exit or from the engine.
The regional view: which exit shows what?
The exit country affects the interface language, the currency format and the ranking of queries that carry a regional signal. But this effect does not show up on every query. On a conceptual question, changing country changes almost nothing, whereas on a query looking for a local service the picture diverges markedly. When building your comparison set, mix these two query types deliberately: one acts as a control, the other as a signal.
For intra-European controls, a German exit is a practical starting point; when a second European reference is needed, a Netherlands exit does the same job with high capacity. To test the Turkish result view, a Turkey exit is used; you can see all the other country options in the location list at the bottom of the page. When changing country, don't forget to carry the browser language preference along with it; otherwise you send the engine two contradictory signals.
One more caution concerns consent screens. When connecting from European Union countries, cookie and data-processing consent screens may appear in different forms; the choices on that screen affect the behaviour of subsequent requests. Make the same choice on every run and write that down too, because sometimes the source of the difference between two runs is not the results but the decision you made on that screen.
DIAGRAMThe role of regional exits in comparison
You can scroll the diagram horizontally to inspect it
The bar lengths represent the relative prominence of the regional signal; it is not an absolute measurement but a suggested order of priority when building your set.
Setup, protocol and leak checks
The protocol choice depends on the nature of the work. If you are working through a browser, an HTTP proxy is the lowest-friction route and it opens a tunnel for HTTPS traffic. Where command-line tools, scripts or non-browser clients are involved, SOCKS5 offers broader compatibility, does not interpret the protocol it carries, and causes fewer problems with unusual clients. In meta search, don't expect a measurable speed difference between the two; base the choice on what the tool you will run the measurement with supports.
Defining the scope at profile level is the safest route for someone running measurements: every request leaving that profile goes through the proxy, and your everyday browser is unaffected. The connection details are entered into the proxy.example.com, 8080, username and password fields; the real values come from your panel. If after setup you are getting 407 , it means credentials are not being sent or your address authorisation has lapsed.
The full list of leak tests and the order in which to run them the verification section on the Swisscows page ; here we add only the part specific to meta search. Before each run, measure that the exit is still up with with the proxy checker tool and write the address you used on that run at the top of your log. On a merged page, a dropped exit and a source that fails to respond produce the same symptom — a shortened results list — and afterwards you can only work out which it was from these two lines.
The second detail specific to meta search is the duration of the run. Because the page waits for the slowest source's response, run times naturally fluctuate; to avoid confusing that fluctuation with instability in the exit, note both the total duration and how many sources contributed to the list on each run. If the duration grows while the source count stays constant, the slowness is on the network side and the exit is where to look; if the source count drops while the duration grows, the problem is in the engine's own chain and changing the exit will fix nothing.
Warning
Do not continue measuring on an exit where you see a certificate warning. A properly set-up tunnel does not interfere with the TLS session; a warning indicates the traffic is being opened and re-encrypted.
Limits, cost and reasonable use
In meta search, most of the latency is structural: the page cannot complete until the slowest source's response arrives. The proxy adds its own hop on top of that time; the two costs do not cancel out, they stack. That is why the page loading later with the proxy on is expected behaviour, and any narrative claiming the proxy shortens latency is technically incorrect. If you want to separate the two, run the same query with the proxy on and off and put the timings side by side: if the difference stays constant from run to run, it is the added hop; if it fluctuates from run to run, the source side is the deciding factor.
On the cost side, search pages produce small responses, so data consumption is rarely the deciding factor. What decides is the number of repetitions and the address requirement. In a study of hundreds of runs, the problem is not bandwidth but concurrency and rate limits; plan accordingly, and write down how many runs, how many queries and how many concurrent connections you will run before you start.
Finally, scope: this page covers regional verification, privacy and research scenarios. A service's terms of use apply regardless of exit country; bulk query generation, fake engagement or defeating security measures are not the subject of this guide. If you want to compare engines of a similar structure, self-hosted SearXNG is a good counterpart; you can find its guide and the setup pages for other engines in the related pages list below.
Frequently asked questions about MetaGer and proxies
01Why do results barely change when I switch the exit country?
On a merged page, the effect originates in the sources' own responses. Because sources that are insensitive to regional signals contribute the same material, the overall picture is smoothed out. If you want to see a difference, use queries with explicit local intent and run the comparison across the whole set.
02Does the proxy affect how the aggregator connects to its sources?
No. Your proxy setting applies only to the connection between you and the search engine. How the engine reaches its own sources lies entirely outside your network configuration and cannot be changed from there.
03Why doesn't exact-phrase search narrow things down as much as I expect?
Even when the operator is passed on to every source, each source interprets it differently. The contribution of a source that supports it narrows, while the list from one that doesn't stays broad, and on the merged page you see the sum of the two. To measure the narrowing, compare against a plain run.
04The results page loads slowly — is the proxy the cause?
Partly, perhaps, but in meta search part of the latency is structural: the page has to wait for the slowest source to respond. The proxy adds its own hop on top of that. To separate the two, run the same query with the proxy off and compare the two timings.
05Do I need to note which record came from which source?
If you are doing comparative work, yes. The source distribution is the most direct clue as to whether the difference between two runs comes from the region or from a shift in source weighting, and it cannot be reconstructed afterwards.
06Do the choices on the consent screen affect the measurement?
Yes. The decision you make on the consent screen can change the behaviour of subsequent requests, and that decision is stored in the browser profile. Make the same choice on every run, write that choice into your log, and if you are going to reset the profile between runs, apply this identically for all exits.
07Is a free exit good enough for this kind of work?
Not on a merged page, because what you lose is not speed but traceability. On a free exit the address can change mid-run, and at that point you cannot match which record came from which source across runs. The ordinary fluctuation in source weighting and the effect of the address change overlap; you are left with no constant to separate the two, and your finding never gets beyond the sentence "the results changed".
08Should I use SOCKS5 or an HTTP proxy?
If you are working from a browser, either will do the job, and HTTP proxy setup is quicker. If you are using a script or a command-line tool, SOCKS5 offers broader compatibility. Base the choice on tool compatibility, and don't expect a difference in speed.