Forum Discussion
Rest api call with multiple pages
Hello,
I need to migrate an SSIS package that make API calls and return multiple pages.
The call goes to https://{{HOSTNAME}}/personnel/v1/person-details, and I can manually do https://{{HOSTNAME}}/personnel/v1/person-details?page=1&per_page=200, my issue is that I do not get any metadata to create the rules on the foreach loop on the first call, and I am not sure how to go about it
I get something like this, no metadata
[
{
"additionalName1": null,
"additionalName2": null,
"nationality1": null,
"nationality2": null,
"nationality3": null,
},
{
"additionalName1": null,
"additionalName2": null,
"nationality1": null,
"nationality2": null,
"nationality3": null,
},
{
"additionalName1": null,
"additionalName2": null,
"nationality1": null,
"nationality2": null,
"nationality3": null,
}
]
good morning, i used at the end a notebook.
something similar to:
import requests import time from requests.adapters import HTTPAdapter from urllib3.util.retry import Retry # API Configuration API_URL = "https://service3.ultipro.ca/personnel/v1/person-details?page={}" USERNAME = "YOUR_USERNAME" PASSWORD = "YOUR_PASSWORD" HEADERS = { "Authorization": "Bearer YOUR_API_TOKEN" } # Configure session with retries session = requests.Session() retry_strategy = Retry( total=5, status_forcelist=[500, 502, 503, 504], backoff_factor=2 # Exponential backoff ) adapter = HTTPAdapter(max_retries=retry_strategy) session.mount("https://", adapter) # Pagination Variables page_number = 1 has_more_data = True all_data = [] while has_more_data: try: response = session.get( API_URL.format(page_number), auth=(USERNAME, PASSWORD), headers=HEADERS, timeout=10 # Set timeout to avoid hanging requests ) if response.status_code == 200: data = response.json() if not data: has_more_data = False break all_data.extend(data) page_number += 1 # Move to the next page elif response.status_code == 429: print("Rate limit hit, retrying after delay...") time.sleep(5) continue else: print(f"Error {response.status_code}: {response.text}") break except requests.exceptions.Timeout: print("Request timed out. Retrying...") time.sleep(5) # Wait before retrying except requests.exceptions.RequestException as e: print(f"Request failed: {e}") break print(f"Total records fetched: {len(all_data)}")
9 Replies
- GilbertQSuper User
Hi AstridM
Here is an example how to do this using power query, which I got from copilot.
let GetPage = (page as number) => let Source = Json.Document(Web.Contents("https://api.example.com/data?page=" & Number.ToText(page))), Data = Source[data], NextPage = if List.IsEmpty(Data) then null else page + 1 in [Data = Data, NextPage = NextPage] in GetPagelet GetAllData = (page as number, accumulated as list) => let Result = GetPage(page), NewAccumulated = List.Combine({accumulated, Result[Data]}), NextPage = Result[NextPage] in if NextPage = null then NewAccumulated else @GetAllData(NextPage, NewAccumulated) in GetAllDatalet AllData = GetAllData(1, {}) in AllData- AstridMAdvocate I
where are you applying all that in the pipeline?
- v-kpoloju-msftCommunity Support
Hi AstridM,
Thank you for reaching out to the Microsoft Fabric Community Forum.
After reviewing the details you provided, here are a few workarounds that might resolve the issue. Please follow the steps below.
Yes, it is possible to migrate an SQL Server Integration Services package that makes API calls and returns multiple pages.-
Please ensure that the API endpoint returns the expected data. Utilize tools like postman or curl to manually test the API calls and verify the response. Confirm that the response includes necessary pagination metadata.
-
APIs supporting pagination include metadata in the response headers or body, such as total_pages, current_page, or next_page. If this metadata is absent, consult the API documentation or contact the API provider for pagination details.
-
If the API lacks pagination metadata, implement custom pagination logic in your SSIS package. Here’s a general approach: Make the initial API call to retrieve the first page of data. Use a ForEach Loop container in SSIS to iterate through the pages, tracking the current page number with a variable.
-
After each API call, check if the response contains data. If the response is empty or contains only null values, assume there are no more pages to fetch.
-
Set up an HTTP Connection Manager to connect to the API endpoint. Use a Script Task to make the API calls and manage the pagination logic, writing the script in C# or VB.NET. Configure the ForEach Loop container to iterate through the pages, updating an SSIS variable with the current page number after each iteration.
Please go through the following link for more information about dynamic parameters:
Foreach Loop Container - SQL Server Integration Services (SSIS) | Microsoft LearnIf this post helps, then please give us Kudos and consider Accept it as a solution to help the other members find it more quickly.
Best Regards.
- AstridMAdvocate I
Hello, I have it working on SSIS, I use a connector called zappy and that is exaclty what I need to do, "After each API call, check if the response contains data. If the response is empty or contains only null values, assume there are no more pages to fetch." my issue is I am not sure how to do it. and all the documentation I found it is pointing me to calls that do return metadata.
- v-kpoloju-msftCommunity Support
Hi AstridM,
Thanks for reaching out to the Microsoft fabric community forum.
After thoroughly reviewing the details you provided, here are few alternative workarounds that might help resolve the issue. Please follow the steps below:
- Create an HTTP Connection Manager in SSIS to connect to the API endpoint. Create an SSIS variable to store the current page number (e.g., Current Page). Create another variable to store the API response (e.g., API Response).
- Use a Script Task to make the initial API call and store the response in the API Response variable. Set the Current Page variable to 1.
- Add a For Each Loop container to iterate through the pages. Configure the loop to continue if there is data in the API Response variable.
- Inside the For Each Loop container, add a Script Task to make the API call for the current page.
- Use the Current Page variable to construct the URL (e.g., https://{{HOSTNAME}}/personnel/v1/person-details?page= + CurrentPage + &per_page=200).
- Parse the response and check if it contains data. If the response is empty or contains only null values, set a flag to exit the loop. After each API call, increment the Current Page variable by 1.
If this post helps, then please give us ‘Kudos’ and consider Accept it as a solution to help the other members find it more quickly.
Best Regards.
-
- AstridMAdvocate I
good morning, i used at the end a notebook.
something similar to:
import requests import time from requests.adapters import HTTPAdapter from urllib3.util.retry import Retry # API Configuration API_URL = "https://service3.ultipro.ca/personnel/v1/person-details?page={}" USERNAME = "YOUR_USERNAME" PASSWORD = "YOUR_PASSWORD" HEADERS = { "Authorization": "Bearer YOUR_API_TOKEN" } # Configure session with retries session = requests.Session() retry_strategy = Retry( total=5, status_forcelist=[500, 502, 503, 504], backoff_factor=2 # Exponential backoff ) adapter = HTTPAdapter(max_retries=retry_strategy) session.mount("https://", adapter) # Pagination Variables page_number = 1 has_more_data = True all_data = [] while has_more_data: try: response = session.get( API_URL.format(page_number), auth=(USERNAME, PASSWORD), headers=HEADERS, timeout=10 # Set timeout to avoid hanging requests ) if response.status_code == 200: data = response.json() if not data: has_more_data = False break all_data.extend(data) page_number += 1 # Move to the next page elif response.status_code == 429: print("Rate limit hit, retrying after delay...") time.sleep(5) continue else: print(f"Error {response.status_code}: {response.text}") break except requests.exceptions.Timeout: print("Request timed out. Retrying...") time.sleep(5) # Wait before retrying except requests.exceptions.RequestException as e: print(f"Request failed: {e}") break print(f"Total records fetched: {len(all_data)}")