Fetching data from APIs using Django and GraphQl without hitting the rate li..

This video features Manaswini Das at DjangoCon Europe 2019 in Copenhagen, Denmark.

Fetching data from APIs using Django and GraphQl without hitting the rate li..
0:18:41
Published April 23, 2019
2,199 views

Fetching data from APIs (GitHub) using Django and GraphQl without hitting the rate limits

https://2019.djangocon.eu/talks/fetching-data-from-apisgithub-using-django-and-gra/

By Manaswini Das: https://twitter.com/ManaswiniDas4

Summary

Manaswini Das explains how she integrated GitHub data into the Open Humans platform using Django, OAuth, GraphQL, and background task processing. She describes rate limits as both a practical constraint and a security measure, then argues that GraphQL can reduce the number of requests by retrieving related data through one precise query and returning a single response. The project authenticates users with Open Humans and GitHub, downloads their GitHub data as JSON, and lets them update it; future work would visualize the data.

Key takeaways

  • GitHub and Twitter were selected as useful data sources for Open Humans because they provide detailed, structured user activity data.
  • Rate limits control API traffic, help maintain service reliability, and reduce the risk of denial-of-service attacks.
  • Celery and requests-respectful were used to manage asynchronous work and respect API request limits.
  • GraphQL provided version stability, precise field selection, and the ability to combine data from multiple sources in one response.
  • The application uses OAuth to authorize Open Humans and GitHub, then stores the user’s GitHub data as a downloadable JSON file.
  • The presenter noted that the next step would be visualizing the imported data.

Summarised automatically from the transcript.

Transcript

2,932 words · auto-generated Show

Automatically transcribed, so expect mistakes in names and technical terms.

0:00

Speaker 1: So uh hello uh world, hello DjangoNots, hello netty zenz. Um yeah I'm here to talk about uh fetching uh data from APIs. Uh using Django and GraphQL without hitting rate limits. It's a pretty long title. It's my first um uh international talk, so I may uh have to work on the titles more. Uh so uh I'm Manishuni Das and I'm a student. from um computer science and engineering from College of Engineering and Technology Bhuneshwar Ordesandya. So this is a project that I developed as a part of my Art Preach internship with Open Humans Foundation and 2018. So Open Humans Platform provides a A single platform for all the

0:47

Speaker 1: users to store their data, including Ancestry DNA, 23andMe Fitbit, all of that. So my project proposal included including um GitHub and Twitter data sources to open humans. So before moving on to that, I would like to uh say who am I? So I'm a student, uh I'm uh almost a graduate, I'm almost there But since I I haven't received my degree yet, so I'm still an undergraduate. Uh I'm from uh I'm also a Processing Foundation Fellow 2019. That's going on now and um a girl script summer of code mentor and an RFiji 2018 alumni as well yeah so uh I had the liberty to choose uh APIs to choose data sources. So uh why did I choose GitHub API and Twitter API is uh due to the fact that uh

1:35

Speaker 1: GitHub API the data is already sorted and organized and which makes it an excellent source for visualization When you see all those green color contributions, uh it gives you an inspiration to work more and to work for that uh darkest green in your uh profile. Um Then we have meticulous information regarding pull requests, issues, activities, repositories. So yeah, that's just a meticulous whole lot of data that you would like to see in your own Open Humans uh account. Um yeah, so moving on. Um so uh I mentioned uh without hitting rate limits. So I need to cover what uh rate limits is all about. So it is like uh the amount of traffic outgoing and then coming to a particular network or a particular API. So uh Yeah, so that's what rate limits are.

2:21

Speaker 1: So uh you may have come across these requests like uh it has already hit the rate limits. GitHub APA has five thousand requests per minute, uh Twitter API has six hundred uh sixty requests per minute, so that is uh rate limits. It ensures better flow of data and increases uh security and uh makes sure uh that You don't uh you know fetch a lot of data and that is not uh visible or not you know not able to organize, uh very difficult to maintain. So Uh yeah, so it increases uh security in the way that it uh prevents DDOS attacks, that is distributed denial of service attack that is not really for uh like uh It doesn't really pose a risk to the se uh security risk, but it renders a website or an online business inaccessible, so it's it's dangerous.

3:09

Speaker 1: Um it prevents users from making an attempt to load on some information. I covered that before, so Yeah. Moving on. These are some of the challenges that I faced uh while working with GitHub API. Uh it uses rate limits, so I had to respect that. And uh take care that if I fetch a lot of data that will not lead to a lot of data, but it will lead to a lot of errors. So uh moving on we have uh fetching data from a GitHub takes time, not because of the rate limits, but it can be a lot of data, uh depending upon your number of countries. contributions. We want to regularly update data and take into account data we already did upload to open humans. Yeah. Open humans has this ability to, you know, um update the data that we already have since Uh you don't uh like uh

3:54

Speaker 1: stop contributing, you just keep adding to your contribution. So yeah, it helps to update your data as well. Uh so uh so for rate limiting I used a salary and requests respectful. Uh for salary it's a simple, flexible, reliable distributed system to provide vast amounts of requests. So uh it is Uh it helped me a real lot and rate limiting. Apart from that, we also have uh we also used a request respectful, which is uh minimalist wrapper uh on requests to you know respect the rate limits. Uh it also helps to scale out a single thread uh single process or even a single machine, it also proxies HTTP requests um for small code changes and It uh works with Python 2 and Python 3 as well. You if you want to check you uh check it out, it is uh the link is in the description like uh github.

4:43

Speaker 1: com slash serpent eye slash request restrictful, you can just uh have a look there Uh so uh I used GraphQL, so it was also upon me. I faced a lot of issues with uh REST APIs while Um REST APIs while fetching data because uh I had to fetch data from a number of APIs, from multiple APIs, so there was a risk of hitting rate limits. Uh so I uh shifted to GraphQL GraphQL is an open source uh server-side runtime specification developed by Facebook. And uh it is uh as the name suggests, it's QL, which is a query language uh to fetch uh uh and mutate, query, fetch and mutate uh data from data sources. is the logo for you. Yeah. Um by GraphQL. Uh yeah. There are some reasons why I changed to GraphQL.

5:29

Speaker 1: Uh the first reason being versioning. Uh REST APIs have a lot of versions. Uh uh taking into account GitHub API, uh REST API's versions are version one, two and three. But uh in GraphQL there is a single version. If there is any updates, there is to the single version itself. So um it made that version uh versioning thing really uh you know handy. Uh apart from that we have uh Ask for what you need. Yeah. If you are asking something and you don't want anything more or less than that, so this is like a very to-the-point and concise concise way to fetch data. Um I don't know whether I can show you a demo here Um

6:15

Speaker 1: let me say. Yeah, I can I think Oops, I can't. Okay, it's fine. Oh uh um Yeah, uh I would like to show the net this screen. Okay, so uh this is the demo. This is the version four of GraphQL API. It has Explorer available to us. Um you can see that uh this is the query on one side and this is uh the response on the other side. Uh whatever you ask for uh like a name with owner Stargazers, sparks, it's just shows the exact data, which is very similar to the query, so you can easily correlate.

7:07

Speaker 1: If you uh the best thing about explorers is that um Uh the thing is uh you can um easily predict what will be your next node or what will be your next edge by simply clicking on control space. Uh so that's the best part about uh GitHub GraphQL API or version 4 Um I don't know whether I can Yeah Um yeah, so this was the demo. Then we have uh get many APIs response in a single request. So that helped me a lot That helped me not writ uh not uh hitting rate limits and um you know stitching a number of uh REST APIs and getting a single response out of that. Because GraphQL provides a schema stitching which is uh really essential for uh

7:52

Speaker 1: querying from multiple REST APIs. This is how GraphQL work works. It has uh the client uh can work and operate via uh you know any computer or any any device. Then it sends a request to the GraphQL, then GraphQL can fetch data from legacy systems, microservices, or third-party APIs and then produce you a single response. Yeah, uh this is OWAT. I know uh all of you know OWAT, right? You have heard about OWAT, I think. How many of you know about OWAT Uh that's pretty a lot of people. But since this talk is for uh intermediate as well as bigners, so I would like to cover what it is. Uh we have almost come across this thing like uh Uh sometimes ESPN

8:38

Speaker 1: asks to, you know, log in with GitHub and then you are uh uh sorry, ESPN asks uh whether you can log in with, you know um Facebook. So you get to log in with Facebook without even entering your password. So that is what OWAuth is all about. So this is how pretty much how OAuth works uh first an authorization request is sent, then um the resource on a sense uh an authorization grant, then the client sends An authorization grant to the authorization server, and then the authorization server uh returns an access token and the access token is then returned to the resource server and then the protected resources sold to the client So that is pretty how pretty much how OWAT works. This is the workflow that I followed in a project.

9:23

Speaker 1: That is first the user goes to the website provided by this repo. Then uh authorizes uh with open humans. You can also authorize with a Google account, so that is another OAuth authentication taking place there. Then we have uh redirect to the GitHub uh page. Uh Then uh you have to authorize with GitHub, that is again login with GitHub. And then uh GitHub data is available as a uh file, as a JSON file. You can easily download that and Yeah, you can alm always click on that update button and get regular updates of your own data that is in JSON format. So demo time, uh you can find uh this work in uh this link. This is uh there. Uh Manishuni Das says uh what Ogithub source, so you can go there and check. Uh so this is the video

10:09

Speaker 1: Uh actually uh the thing is um I was trying to uh run it in my own computer but uh there was some issue and it suddenly stopped working. You you can relate to me, right? Uh yeah, so uh so I ha I somewhat managed to uh take uh you know a screen recording of that so I'd like to play it Oh, sorry. Um, I think it's not plain

10:55

Speaker 1: Yeah, so first um first you have to uh start the server which is uh by me by typing Hiroku local in your terminal or going to manage. py run server and then you have uh this app running at localhost 5000 Yeah, so just go on to localhost 5000. Then you have this home screen which helps you integrate your GitHub account to Open Humans and this is the homepage. So you have to click on start with open humans authorization Um yeah, uh then it directly leads to this page, but actually it uh skipped the award part. You can sign in and sign up at the you know first thing. Then uh yeah, so this is authorized GitHub data upload. Then you have to authorize this app

11:42

Speaker 1: by clicking on authorize project. Then Out. Yeah, so uh this leads to the uh redirect page that is uh for logging in into GitHub and then yeah, you have to log in into GitHub. At first, your data may not be visible, uh, but uh when you reload this, your data will be visible. Yeah, you can all see your data and uh in the terminal as well, like this. How load of data? Uh it's in JSON format, so Yeah, uh

12:27

Speaker 1: just reload it and we get download GitHub data. So uh when you click on download, you have that file available in your local system. Yeah, so again that's uh pretty lot of data. All the data will be as available. Yeah Say uh so that's how I fetched data from uh uh from GitHub API and made it visible to the Open Humans platform. This data will be visible in your open humans account as well. I think it's still loading.

13:20

Speaker 1: Um it's boring. Uh okay, we'll we'll switch to the next slide. Yeah. So uh this was all I had to say. Um uh very uh uh I extend my heartfelt gratitude to DjangoCon Europe for organizing this awesome conference and having me speak uh my work out and um

14:05

Speaker 1: I also uh have to uh give a huge shout out to Outreachy for giving me an opportunity to work with Open Humans Organization. I would like also like to thank my mentor Mike Escalante, Bastian, and Matt Price Ball for extending me the help and support. And um apart from that, I'm very grateful to my parents and my friends for uh helping me reach where I'm today. So uh that's all I had to say. Thank you.

14:38

Speaker 2: That was awesome. Thank you so much. So we have a few minutes for uh some questions. Uh again, we are using DjangoCon QA online and you can also um come up present. Um in in the center, if you are in the auditorium.

14:59

Speaker 1: Any questions? I know it's already lunchtime, but yeah, we can answer questions. Oh ,

15:06

Speaker 3: I have one question. Um Did you use the actual was it only a a test or did you actually use the GitHub data? Uh with anyone?

15:17

Speaker 1: I did use GitHub data.

15:19

Speaker 3: For what?

15:21

Speaker 1: For?

15:22

Speaker 3: For what? Like why f no uh I mean what what did you do with it? Like what what's it?

15:27

Speaker 1: It just gets uploaded to the Open Humans platform and your data is visible like uh along with the other data sources that Open Humans provides, such as twenty threeandme. um and uh ancestry DNA and uh all of that. It uh the next future scope is uh you know visualizing your data. That's it.

15:44

Speaker 3: Thank you very much.

15:45

Speaker 1: Thank you.

15:47

Speaker 4: Hi, um so I've got a question for you in terms of the the practicalities of building a system like this. Okay. Um GraphQL is well documented and they've got its APIs. OpenAP uh OAuth is documented, has its APIs, but the actual mechanics of trying to get one to talk to the other and making all the pieces fit together can be complicated Do you have any tips or suggestions for someone trying to debug why their big system like this isn't working? Any tools or tips or tricks that were useful to you building this system?

16:18

Speaker 1: Okay, uh uh a lot of my friends reached out to me. I used uh like Stack Overflow was my home that time 'Cause uh I had to deal with a lot of uh you know errors. There was a simple error, like all of the that stuff was working all good and fine, but uh suddenly the data wasn't sewing. So uh I just had to, you know, uh there was this tags, uh so I just had to give uh GitHub tag. So that's uh that's what uh that's all I had to do. There may be minor errors, there may be major errors, but uh it's better if you reach out and you know to forums. I don't know whether that answer your questions or not. Thank you.

16:55

Speaker 5: Hi, thank you for your talk.

16:56

Speaker 1: Hi.

16:57

Speaker 5: Um I had a question about uh you mentioned for as part of the Open Human Project uh you need to with a certain number of requests that you can make before you get rate limited, you need to uh figure out uh timely updates versus uh bringing in new records. So if you're only working with a thousand if you have millions of records but you only have a thousand API calls per minute, how do you decide how many of those are going to be uh bringing in new records versus how many uh trying to check and confirm that your historical records are still up to date?

17:28

Speaker 1: Um I think Gras RafQL uh provides everything in a single API call. So uh so that pretty much worked for me, but I don't know whether uh uh I didn't have a chance to explore the limits. Uh thanks for the question. Okay, cool.

17:45

Speaker 6: Uh there is a question from the internet. If there are any high-profile GraphQL API providers apart from GitHub Like are there any gre uh most APIs tend to be REST APIs or XML stuff? Um if there are any other uh GraphQL APIs that you're aware of that are rather large apart from GitHub.

18:10

Speaker 1: I think there's a Twitter thing as well. I was really lucky to find a GraphQL cohort for uh the REST API of GitHub. So I think it's pretty m uh pretty much difficult if there is no documentation for uh GraphQL APIs of different data sources, uh, but I was lucky to find the GitHub uh cohort.

18:27

Speaker 6: Thank you.

18:28

Speaker 1: Thank you. Any other questions?

18:35

Speaker 2: That was awesome. Thank you so much.

Questions this talk answers

What are API rate limits, and why do APIs use them?

Rate limits restrict how much traffic or how many requests can be sent to an API. They help maintain reliable data flow, improve security, and reduce the risk of abuse such as denial-of-service attacks.

Discussed at 1:35

How can I avoid hitting API rate limits when fetching data with Django?

The project used Celery for distributed request handling and `requests-respectful` to make HTTP requests while respecting rate limits. It also switched from several REST requests to GraphQL, which can return data from multiple APIs in one request.

Discussed at 3:53

Why use GraphQL instead of REST APIs for fetching data?

GraphQL avoids maintaining multiple API versions and lets clients request exactly the fields they need. Its schema stitching and single-request model also make it useful when combining data from multiple sources.

Discussed at 5:29

How does GraphQL reduce the number of API requests?

A client sends one GraphQL query, and GraphQL gathers the requested data from legacy systems, microservices, or third-party APIs before returning one response. This avoids stitching together many separate REST responses in the client.

Discussed at 7:07

How does OAuth work when connecting GitHub data to Open Humans?

The user authorizes the application with Open Humans and then authorizes it with GitHub. The application receives access through OAuth and makes the GitHub data available as a downloadable JSON file that can be updated later.

Discussed at 9:23

How do I upload GitHub data to Open Humans using the demo project?

Run the application locally, open its localhost page, start Open Humans authorization, authorize the project, and then log in to GitHub and approve access. After reloading, the GitHub data appears and can be downloaded as a JSON file.

Discussed at 10:55

What tools can help debug a Django, GraphQL, and OAuth integration?

The speaker relied heavily on Stack Overflow and recommends reaching out to developer forums when the integration fails, including for small issues such as missing tags or data not appearing.

Discussed at 16:18

How can I handle both new records and updates without exceeding an API request limit?

The speaker’s approach was to use GraphQL to obtain the needed data in a single API call, though she noted that she had not explored the exact limits in depth.

Discussed at 17:28

Are there large GraphQL APIs besides GitHub?

The speaker mentioned a Twitter GraphQL offering, but said that finding GraphQL versions of other data sources can be difficult when they lack documentation. She considered herself fortunate to find a GraphQL counterpart for GitHub’s REST API.

Discussed at 18:10

Presenters

Note: We understand that names change, people change, and bodies change. We respect each individual's journey and privacy. If you have any concerns about a video or need us to remove content, please don't hesitate to contact us. We will handle your request with care and promptly address any issues.

More videos from DjangoCon Europe