Advanced Web Programming

Web Architecture, Internet Protocols and HTTP

PGCP-AC

Every web application depends on an exchange between a client and a server. Understanding that exchange makes AJAX, REST APIs, authentication, Express, React and deployment easier to reason about.

The Internet and the Web are often used as if they were the same thing, but they describe different layers. The Internet is the network infrastructure that moves data between connected systems. The Web is one service built on that infrastructure. A browser uses web protocols to request resources, but the same Internet also carries email, secure shell sessions, voice calls, file transfers, database traffic and many other services.

1. Internet, Web and application architecture

The Internet is a global system of interconnected networks. Internet Protocol provides addressing and packet delivery between machines. Email, file transfer, voice communication and the Web all use this infrastructure.

The World Wide Web is a system of addressable resources exchanged primarily through HTTP. A resource may be an HTML page, image, stylesheet, JSON document, video or an operation exposed by an API.

LayerResponsibilityExample
ClientPresents the interface and initiates requestsBrowser or mobile app
Web serverAccepts connections and serves static resourcesApache or Nginx
ApplicationApplies business rules and constructs responsesNode.js/Express service
DatabaseStores persistent informationMySQL or MongoDB

These roles may run on one machine during development and on many machines in production.

2. Development of the Internet and the Web

Early computer networks connected a small number of systems for resource sharing and research. Packet switching became a central idea: instead of reserving one complete physical circuit for an entire conversation, a message is divided into packets that can share network links with other traffic. Each packet carries addressing and control information and the receiving system reconstructs the original data.

ARPANET demonstrated wide-area packet networking. The adoption of the TCP/IP protocol family allowed independently designed networks to communicate as an internetwork. Domain names later made hosts easier to identify than numeric addresses alone. The Internet expanded from research and institutional networks to commercial providers, homes, mobile devices, cloud data centres and embedded systems.

The World Wide Web was developed as a linked information system using three foundational ideas:

  • a uniform addressing system for resources, now expressed through URIs and URLs;
  • HTTP for exchanging requests and responses;
  • HTML for representing linked documents.

The earliest pages were mainly static documents. Server-side programming introduced generated pages and database-backed applications. JavaScript made browsers programmable, while AJAX enabled background requests without complete page reloads. Modern systems include single-page applications, APIs, streaming, real-time connections, content delivery networks, cloud platforms and distributed services. These additions still depend on naming, addressing, routing, transport, security and application protocols.

3. Packet delivery across networks

When an application sends data, the networking software divides responsibility into layers. Each layer adds control information needed by its peer at the destination. A simplified TCP/IP view is:

LayerMain responsibilityExamples
ApplicationMeaning and format of exchanged messagesHTTP, DNS, SMTP, SSH
TransportCommunication between processesTCP, UDP, QUIC
InternetAddressing and routing between networksIPv4, IPv6, ICMP
LinkDelivery across one local network segmentEthernet, Wi-Fi

Suppose a browser requests a page from another country. The HTTP message becomes transport data, which is placed into IP packets, which are carried inside link-layer frames on each local hop. A home router forwards packets to an Internet service provider. Other routers examine destination network information and select successive paths. The link-layer frame changes at each hop, while the destination IP address normally continues to identify the remote endpoint.

Routing is distributed. A packet does not contain a complete guaranteed physical route and different packets may take different paths. Routers use routing tables built from directly connected networks, configured routes and routing protocols. Congestion, failures and policy can change the selected path.

Packet delivery can fail, arrive late, arrive more than once or arrive out of order. The Internet Protocol itself provides best-effort delivery rather than an application-level guarantee. Higher layers decide whether reliability, ordering, retransmission and flow control are required.

4. TCP, UDP, ports and sockets

An IP address identifies a network interface or logical endpoint at the Internet layer. A port identifies a receiving process or service at the transport layer. The combination of protocol, address and port helps distinguish concurrent communication.

TCP is connection-oriented. Before application data is exchanged, the endpoints establish connection state. TCP provides an ordered byte stream, retransmits missing data, detects duplicates and regulates sending through flow and congestion control. HTTP/1.1 and HTTP/2 commonly operate over TCP.

UDP sends independent datagrams without establishing TCP-style connection state. It does not itself guarantee delivery, order or duplicate suppression. Its small transport overhead suits cases in which the application supplies its own policy or values timeliness over retransmission. DNS commonly uses UDP for ordinary queries and can use TCP when required.

QUIC builds secure, reliable, multiplexed transport behaviour over UDP and is used by HTTP/3. Implementing reliability above UDP does not make QUIC unreliable; it means the reliability mechanism belongs to QUIC rather than TCP.

Well-known ports provide conventional service locations. HTTP commonly uses port 80 and HTTPS commonly uses port 443. A server can listen on another port when the URL or configuration identifies it. A socket is a programming abstraction for one endpoint of network communication. A listening server socket accepts client connections and creates connected communication endpoints.

5. IP addressing and local delivery

IPv4 addresses contain 32 bits and are commonly written as four decimal octets, such as 192.0.2.10. IPv6 addresses contain 128 bits and use hexadecimal groups, providing a much larger address space and other protocol improvements.

A subnet prefix identifies which address bits represent the network. Before sending, a host determines whether the destination is on its local network. For a local IPv4 destination, Address Resolution Protocol discovers the link-layer address associated with the IP address. For a remote destination, the host sends the frame to a configured gateway, which routes the packet onward. IPv6 uses Neighbor Discovery for related local functions.

Private IPv4 address ranges are used inside many local networks and are not routed as ordinary public Internet addresses. Network Address Translation at a router can map several private endpoints to public communication state. NAT conserves public IPv4 addresses and hides internal addressing structure, but it is not a replacement for a firewall or application authorization.

An address such as 127.0.0.1 refers to the local IPv4 loopback interface. localhost commonly resolves to a loopback address and lets software communicate with services on the same machine without traversing an external network.

6. Domain names and DNS resolution

People remember names such as www.example.com more easily than network addresses. The Domain Name System is a distributed, hierarchical database that maps names to records.

The name is read hierarchically from right to left:

www.shop.example.com
│    │    │       └─ top-level domain: com
│    │    └──────── registered domain: example.com
│    └───────────── subdomain: shop
└────────────────── host or service label: www

Important record types include:

RecordPurpose
AMaps a name to an IPv4 address
AAAAMaps a name to an IPv6 address
CNAMECreates an alias to another canonical name
MXIdentifies mail exchangers for a domain
NSIdentifies authoritative name servers
TXTStores text used for verification and policy information

A typical resolution proceeds as follows:

  1. The browser and operating system check their caches and local configuration.
  2. A stub resolver asks a configured recursive resolver.
  3. If the answer is not cached, the resolver follows referrals from root servers to top-level-domain servers and then authoritative servers.
  4. The authoritative answer is returned and cached according to its time to live.

The recursive resolver performs work for the client. An authoritative server publishes records for a zone. DNS caching reduces latency and load, but it means record changes are not visible everywhere immediately. DNS supplies addressing information; it does not by itself prove that the server is trustworthy. HTTPS certificate validation provides server authentication for the requested name.

7. Web architecture and server roles

Web architecture is based on clients requesting representations or operations from servers. A browser is a user agent: it resolves names, establishes connections, sends requests, enforces browser security rules, parses HTML and CSS, executes JavaScript and renders the interface.

A web server accepts HTTP connections. For a static resource, it may read a file and return it directly. For dynamic work, it can forward or proxy the request to an application process. The application validates input, applies business rules, accesses databases or other services and constructs a response.

Common server products include Apache HTTP Server, Microsoft Internet Information Services and Nginx. Their capabilities include static-file delivery, virtual hosting, TLS termination, logging, compression, caching, access control and reverse proxying. Node.js can also create an HTTP server directly, although production deployments often place a reverse proxy or managed gateway in front of an application.

A forward proxy acts on behalf of clients, while a reverse proxy acts as an entry point for servers. A reverse proxy can distribute requests across application instances, terminate TLS, cache suitable responses, apply limits and hide internal addresses. A load balancer selects a healthy backend according to an algorithm and operational policy.

A content delivery network places cached content at geographically distributed edge locations. A user may connect to a nearby edge rather than the origin server, reducing latency and origin load. Dynamic or personalized content still requires careful cache rules.

8. URL anatomy

Consider this URL:

https://shop.example.com:8443/books/42?format=pdf#reviews
ComponentValuePurpose
SchemehttpsSelects protocol and security rules
Hostshop.example.comIdentifies the server by name
Port8443Identifies the receiving network service
Path/books/42Identifies a resource within the origin
Queryformat=pdfSupplies additional request parameters
FragmentreviewsIdentifies a position in the representation

The fragment is interpreted by the client and is not sent in the HTTP request target. An origin consists of scheme, host and port. Therefore, https://example.com and http://example.com are different origins, as are ports 443 and 8443.

9. From URL to displayed page

A simplified browser exchange follows these steps:

  1. The browser parses the URL and checks relevant caches.
  2. DNS resolves the host name to an IP address when no usable answer is cached.
  3. The client establishes a transport connection. HTTPS also performs a TLS handshake.
  4. The browser sends an HTTP request.
  5. The server routes it and may call application and database code.
  6. The server returns an HTTP response.
  7. The browser processes the representation and may request additional resources.

Connection reuse, proxies, content delivery networks and service workers can alter the path, but the request–response model remains central.

10. HTTP message structure

An HTTP message contains a start line, header fields, a blank line and an optional body.

GET /api/books/42 HTTP/1.1
Host: example.com
Accept: application/json

The method states the operation, the target identifies the resource and Accept states which response formats the client can process.

HTTP/1.1 200 OK
Content-Type: application/json
Cache-Control: max-age=60

{"id":42,"title":"Web Systems"}

Content-Type describes the enclosed representation. This differs from Accept: the former describes the body that is present, while the latter describes the response formats desired by the client.

11. HTTP methods

MethodNormal purposeSafeIdempotent
GETRetrieve a representationYesYes
HEADRetrieve metadata without response contentYesYes
POSTSubmit data for processing or create under a collectionNoNot generally
PUTCreate or replace state at a known targetNoYes
PATCHApply a partial modificationNoNot guaranteed
DELETERequest removal of a resourceNoYes
OPTIONSDiscover communication optionsYesYes

A method is safe when its defined meaning is essentially read-only. Logging a GET does not make it unsafe because logging is incidental rather than the requested action.

A method is idempotent when repeating an identical request has the same intended effect as performing it once. Repeating DELETE /books/42 may return 204 first and 404 later, yet the intended final state—book 42 absent—remains the same. Idempotence does not require identical responses.

Repeating a POST may create two records. A client should therefore avoid automatically retrying a non-idempotent operation unless the application provides protection such as an idempotency key.

12. Status codes

  • 1xx: informational;
  • 2xx: successful handling;
  • 3xx: redirection or cache-related handling;
  • 4xx: a problem with the request or permission;
  • 5xx: a server failure while handling an apparently valid request.
CodeMeaningTypical use
200OKSuccessful retrieval or update with a body
201CreatedNew resource created, often with Location
204No ContentSuccess without a response body
400Bad RequestMalformed syntax or invalid input
401UnauthorizedAuthentication is absent or unacceptable
403ForbiddenThe server refuses the operation
404Not FoundThe target resource was not found
409ConflictRequest conflicts with current resource state
500Internal Server ErrorUnexpected server failure

Returning 200 with an error object for every failure prevents clients, caches and monitoring systems from interpreting results correctly.

13. Statelessness, cookies and sessions

HTTP is stateless: a request does not automatically remember an earlier request. Applications add continuity explicitly.

Set-Cookie: sessionId=abc123; Secure; HttpOnly; SameSite=Lax

The browser may later return this cookie. The cookie is client-held data; the corresponding session record is often server-held data. They are related but are not the same object.

  • Secure restricts transmission to secure connections.
  • HttpOnly prevents ordinary client-side JavaScript from reading the cookie.
  • SameSite controls when it accompanies cross-site requests.

Session identifiers should be unpredictable, protected in transit and invalidated when the authenticated session ends.

14. HTTPS and TLS

HTTPS is HTTP carried over TLS. TLS supplies encryption in transit, integrity protection and server authentication through certificates. It does not validate user input, prevent SQL injection, correct broken authorization or make unsafe application logic secure.

During a TLS handshake, the client and server agree on cryptographic parameters, establish shared keys and authenticate the server through its certificate chain. A certificate binds a public key to one or more domain names. The client checks the requested host name, validity period, signatures and chain to a trusted certificate authority.

After the handshake, efficient symmetric keys protect application data. Integrity protection detects modification in transit. Forward-secret key exchange can prevent a later compromise of the server's long-term private key from decrypting earlier recorded sessions.

Mixed content occurs when an HTTPS document loads an insecure HTTP resource. Browsers restrict these loads because an insecure script or frame could undermine the protected page. Applications should redirect HTTP to HTTPS, mark authentication cookies Secure and HttpOnly, avoid placing secrets in URLs and deploy Strict-Transport-Security only after all required subdomains and resources support HTTPS.

15. HTTP versions and connection management

HTTP has evolved while preserving the request, response, method, status, header and representation model.

VersionTransport formConnection behaviourMain characteristics
HTTP/1.0Textual messages, normally over TCPCommonly one request per connectionSimple, but repeated connection setup is costly
HTTP/1.1Textual messages over TCPPersistent connections by defaultMandatory Host, chunked transfer, improved caching
HTTP/2Binary framing over TCPMultiplexed streams on one connectionHeader compression and concurrent stream progress
HTTP/3HTTP semantics over QUIC and UDPMultiplexed QUIC streamsIntegrated modern security and reduced cross-stream transport blocking

Persistent connections reuse transport and security setup for several requests. This reduces handshake cost and latency. Servers still apply idle timeouts and request limits so unused connections do not consume resources indefinitely.

HTTP/1.1 pipelining permits later requests before earlier responses arrive, but responses remain ordered and practical deployment has been difficult. Browsers commonly used several connections instead. HTTP/2 divides messages into binary frames and multiplexes streams, allowing several responses to progress over one connection. Because all streams share TCP, a lost TCP packet delays delivery until TCP repairs that part of the byte stream.

HTTP/3 uses independent QUIC streams, so loss affecting one stream does not impose the same transport-level delay on unrelated streams. These changes improve delivery but do not alter the meaning of GET, POST, status codes, caching, authentication or authorization.

Protocol negotiation lets a client and server select a mutually supported version. The https scheme describes security and default-port conventions; it does not by itself specify HTTP/1.1, HTTP/2 or HTTP/3.

16. Caching and validation

Caching reduces latency and server work by reusing a stored response when permitted. Cache-Control expresses freshness and storage rules. A validator such as ETag enables a conditional request. If the stored representation is still current, the server can return 304 Not Modified without sending its body again.

Freshness determines how long a stored response can be reused without contacting the origin. max-age supplies a freshness lifetime. no-store asks caches not to store the response. no-cache permits storage but requires successful validation before reuse; it does not mean “never store.” private limits reuse by shared caches, while public permits shared caching when the response is otherwise suitable.

An ETag is an opaque validator chosen by the server. A client can send it in If-None-Match. If the selected representation has not changed, the server returns 304 Not Modified without the full response body. Last-Modified and If-Modified-Since provide time-based validation, although timestamps may be less precise.

The Vary header identifies request headers that influence representation selection. For example, varying by Accept-Encoding prevents a cache from incorrectly serving a compressed form to a client that did not request it. Personalized or sensitive responses require careful directives. Cache behaviour depends on the method, status, response headers, validators, authentication context and each participating browser or intermediary.

17. Complete request trace

POST /api/books HTTP/1.1
Host: library.example
Content-Type: application/json
Accept: application/json

{"title":"HTTP Fundamentals"}

A suitable response is:

HTTP/1.1 201 Created
Location: /api/books/73
Content-Type: application/json

{"id":73,"title":"HTTP Fundamentals"}

POST requests creation under the collection. Content-Type identifies the submitted JSON, 201 reports creation and Location identifies the newly created resource.

The complete exchange includes several layers:

  1. The browser resolves library.example through DNS.
  2. It selects an address and opens the required TCP or QUIC transport.
  3. For HTTPS, TLS authenticates the server and establishes protected keys.
  4. The client sends headers and the JSON representation.
  5. A reverse proxy may terminate TLS and route the request to an application instance.
  6. The application authenticates the caller, validates input, checks authorization, applies business rules and commits data.
  7. The response crosses the proxy and network back to the browser.
  8. The browser checks the status and Content-Type before interpreting JSON.
  9. Connection reuse and cache metadata influence later exchanges.

A failure at each step has a different meaning. DNS failure means the name could not be resolved. Connection refusal means no reachable service accepted the endpoint. TLS failure means secure establishment or server authentication failed. An HTTP 404 is different: networking succeeded and the server deliberately returned an application-protocol result.

18. Practical considerations

  1. The Internet and Web are related but are not synonyms.
  2. A URL fragment is not sent to the server.
  3. Safe and idempotent describe different properties. DELETE is idempotent but unsafe.
  4. Idempotence concerns intended effect, not identical response messages.
  5. HTTPS protects transport; it does not repair application vulnerabilities.
  6. A cookie and a server-side session are not the same object.
  7. 401 concerns authentication, whereas 403 expresses refusal.
  8. Accept describes a desired response; Content-Type describes an enclosed body.
  9. GET must not be used for destructive actions because crawlers and prefetchers may follow it.
  10. DNS maps names to records; application routing maps an HTTP target to a resource or handler.
  11. A port identifies a transport service, whereas a URL path identifies an application resource.
  12. HTTP/2 changes framing and connection use without changing HTTP method semantics.
  13. no-cache requires validation before reuse; no-store prevents storage.
  14. A reverse proxy represents servers to clients, while a forward proxy represents clients to servers.

Continue learning

Related notes

Put this topic into timed practice

Open mock tests when you want full-exam pacing, or keep drilling in practice mode.